Conceptio › Archive › arXiv CS
arXiv CSopen access

The OCUDU dApp Platform: An Open Runtime and E3 Interface for Real-Time AI-RAN

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

The OCUDU dApp Platform: An Open Runtime and E3 Interface for Real-Time AI-RAN Timothy O’Shea, Matthew Pennybacker, Andriy Kharchenko

arXiv:2609.07843v1 [cs.NI] 7 Sep 2026

DeepSig Inc., Arlington, VA, USA

Abstract—Machine learning has shown its largest gains in the band below 10 ms inside a 3GPP new radio (NR) 5G distributed unit (DU): link adaptation, per-slot scheduling, channel estimation, and the receiver itself. No open platform has let independently built software run there. Prior dApp frameworks reached the band only as external observers of an export stream. This paper is a guided introduction to the OCUDU dApp platform, an open runtime and E3 interface under which signed AI-RAN applications execute inside a production DU under three timing contracts: resident on the GPU receive chain (Class A), inside the scheduler’s 100 µs admitted deadline (Class B), or as never-blocking observers whose results the scheduler consumes (Class C). The conventional path is never displaced, and every authority is typed, validated, and operator-bounded. The paper explains how the runtime, the embedded E3 agent, and the three public repositories fit together; shows a dApp’s source, its signed package, and its lifecycle state machine; defines the contracts a module is written against; and shows how one management surface serves a Python script, an operator’s console, and an LLM agent. On a GB10 gNB with attached handsets, dApps of all three classes, including an out-of-tree neural equalizer, ran together on a live cell without a single fallback, and equalizer variants were compared over the air by lifecycle operations alone. Every measured checkpoint is reported with its conditions and its gaps. Platform, SDK, and a zero-hardware quickstart are public under BSD-3-Clause-Clear as a preview release of the OCUDU AI-RAN Working Group 2, inviting feedback, new use cases, and independent vetting ahead of upstreaming into the OCUDU mainline. Index Terms—AI-RAN, dApp, O-RAN, E3, 5G NR, DU, GPU, real-time control, ISAC, spectrum sensing, neural receiver, MLOps

I. I NTRODUCTION The O-RAN control architecture leaves its most valuable real estate unmanaged. rApps operate at 1 s and above, and xApps reach down to roughly 10 ms [1], [2]. Everything faster has been sealed inside the vendor’s DU: link adaptation, perslot scheduling, channel estimation, and any reaction to the spectrum inside a coherence time. Yet this band is where learned components have shown the largest gains, from neural receivers [3] to learned link adaptation, and it is where sensing applications such as integrated sensing and communication (ISAC) must live to see the channel at all. The dApp concept named this tier and proved the demand. Bonati et al. proposed it [4]; Lacava et al. gave it the E3 interface, the real-time interface between a dApp and its host RAN node, and a working OpenAirInterface framework [5]– [7]; NVIDIA’s Aerial dApp container echoed it [8]–[10]. The O-RAN research arm has since published dApp use cases [11], vendors are opening their lower PHY [12], and deployment

characterizations are appearing [13]. These systems share one structural ceiling, shown hatched in Fig. 1: the dApp is an external observer loop. L1 exports IQ or KPIs, a separate process consumes them, and control returns as a coarse indication with no admission bound, no deadline, and no fallback. The highest-value applications, a neural receiver on the L1’s own tensors or a learned policy answering before the scheduling decision commits, cannot be expressed that way. This paper introduces the OCUDU dApp platform, which opens the sub-10 ms band as a managed, open runtime. The platform rests on two claims. First, AI-RAN’s upside arrives quickly only if this band is open to a diverse community of vendors, operators, and researchers. Second, one infrastructure can carry everything that community builds. The same package format, E3 plane, and safety machinery that let a Python script watch CRC verdicts also let a neural receiver replace the receive chain, a sensing worker steer the scheduler, a continuallearning loop run from live capture to model rollback, and several vendors ship dApps that one operator composes at runtime. The paper is written as a tutorial. It proceeds from the architecture at a glance and the vocabulary it uses (Section III) through the three execution classes and their contracts (Section IV), packages, trust, and the lifecycle state machine (Section V), the E3 plane (Section VI), where the code lives (Section VII), a first session with real source for a Class B module, a portable Class C client, and an agent (Section VIII), the applications it carries (Section IX), the vendor ecosystem (Section X), the latency and safety architecture with its fallback semantics (Section XI), the measured evidence with its limits (Section XII), and what the expanded contract buys (Section XIII). The measured interface study behind the class definitions is reported separately [14]. II. P RIOR DA PP F RAMEWORKS AND T HEIR B OUNDARY The reference dApp implementation [5] pairs a Python authoring library with libe3 [7], a patched OpenAirInterface gNB exporting IQ through an E3 service model, and a FlexRIC extension for xApp–dApp coordination; its canonical dApp senses the spectrum on exported IQ and returns a PRB blacklist for the MAC scheduler to enforce. NVIDIA’s container [8], [9] pairs an E3 manager with a Triton inference server beside the CUDA L1, publishing RAN data through sharedmemory descriptors to GPU inference backends, applied to real-time uplink interference detection in InterfO-RAN [15].

below the RIC’s reach: no standardized tier before dApps

dApps — co-located with the DU

rApps · Non-RT

xApps · Near-RT RIC

1 slot = 500 µs (30 kHz SCS)

1 µs

10 µs

100 µs

1 ms

10 ms

100 ms

1s

10 s

control-loop latency prior dApp frameworks: external observer loop, IQ export + coarse indications

Class A · inline L1

Class C · async

Class B

slot-synchronous, 100 µs GPU-resident replacement bounded intents

observe + advise; producer never waits; closes the loop

OCUDU dApp platform: one package format, one E3 plane, three timing contracts Fig. 1. The control-loop ladder. O-RAN’s RICs stop at about 10 ms. Prior dApp frameworks entered the sub-10 ms band only as external observer loops (hatched). OCUDU covers the whole band with three timing contracts under one package format and one E3 plane.

Both are genuine sub-RIC control, and both validated E3 as the interface. Three consequences follow from the external-loop shape regardless of implementation quality. First, inline applications cannot be expressed: a neural receiver must run on the L1’s tensors and stream inside the slot, and an external process pays serialization, wakeup, and CUDA-context costs to reach data the PHY already holds [16]. Second, the radio inherits the dApp’s failure modes: an experimental component can degrade or stall the cell, because there is no producer-side isolation, admission control, deadline, or always-available conventional path. Third, operations do not scale: every dApp is a bespoke deployment, because there is no signed packaging, typed configuration, model lifecycle, or uniform observability, and no framework gives an operator the machinery to compose dApps from several vendors on one cell. The platform addresses each while keeping the observer pattern fully supported as its entry point (Class C over E3), so existing dApp-style consumers port directly. III. T HE P LATFORM AT A G LANCE OCUDU, the Linux Foundation’s open-source 5G CU/DU project [17], is a production-grade 5G NR stack whose GPUresident CUDA L1 and fronthaul acceleration are described in [16]. The dApp runtime adds three seams to that stack and one management plane over them (Fig. 2). Table I states the three seams side by side, and Table II defines the terms the rest of the paper relies on. The terms are short, but several of them, shadow, evaluation mode, generation, and lane, carry precise meanings that the prose below depends on. A dApp is a shared object built in the vendor’s own repository against the installed SDK, never against the platform’s source tree. It is described by a signed manifest and an SPDX software bill of materials (SBOM), staged into a catalog named in the gNB YAML, and driven through an explicit lifecycle over E3 (Section V). The third-party boundary is a frozen, size-tagged C ABI with a checked-in layout fingerprint

verified across GCC and Clang, C11 and C++17, loaded with dlopen [18]. Cooperation flows only through host-owned, schema-versioned caches. The contracts are backend-neutral: a package declares a CPU (x86, ARM) or CUDA backend and receives matching resident hooks, and the SDK ships every reference in both variants. The emphasis is nonetheless GPUfirst, because the compute headroom beside the baseband lets AI-RAN applications be tried at full fidelity, then optimized, without displacing the radio’s own budget. Four invariants hold everywhere: the host owns memory, streams, validation, and fallback; the conventional path is never displaced; every authority is typed, validated field by field, and operatorbounded; and dApps never call each other. IV. T HE T HREE E XECUTION C LASSES A. Class A: resident inline L1 The host supplies GPU-resident tensors and its own perlane CUDA stream [19], [20]. The input is the full-slot device grid in complex BF16, an explicit DM-RS pilot tensor with bounded coordinates, compact ordered data-RE indices, and typed PUSCH metadata (allocation shape, ports, layers, modulation, noise floor, DM-RS layout), so a module can estimate a channel without calling into the private PHY. The outputs are caller-owned device tensors the host prepares and validates: port-and-layer-major channel estimates plus noise at the first depth; compact equalized symbols in [data_re, layer] order plus post-equalization noise at the second; FP16 soft bits in [data_re, layer, bit] order before descrambling at the third. The three depths rejoin the chain at successively deeper points, and no hook replaces equalization alone, because the equalizer interface also implements channel estimation. A module that succeeds at the second or third depth also supplies the scheduler’s uplink SINR from its own postequalization noise, because the conventional measurement kernels are skipped on that path. Before any module code runs, the host checks the grant against the interface’s declared admission profile with a fixed

OCUDU DU process (gNB): frozen, size-tagged C ABI v1 E3 clients: any process, any host

Class A module CE, CE+EQ, or full RX on the lane’s CUDA stream; completion event invoke, shape admitted

UL grid (GPU)

conventional stage armed as fallback

portable Class C dApps any language, container, or host

DU-owned outputs

EQ

CE

demap

recorders, dashboards

E3AP APER / SCTP 36423

LDPC

ocudu_e3 Python client

dappctl operations

MCP server LLM agents

E3DP FlatBuffers / SCTP 38472 capture candidates, features

Class B module direct call, 100 µs deadline, allow flags

MAC scheduler decision boundary

validated intents

slot publisher try-lock + eventfd

lease

Class C native leased view, 0 copies

host-owned caches interference map, quiet reservations

credentialed local socket

worker process, same host sharedmemory ring or CUDAIPC pool

results

Class C supervised worker seccomp, heartbeats, restart

1 copy, 1 eventfd; full ring = counted drop manages every class off the hot path

embedded E3 agent lifecycle, config, models, metrics, incidents, streams, RAN control

telemetry publisher 8 streams, rate floors, idle = 1 atomic load

supervise, restart

SEQPACKET control credentials checked

Fig. 2. The platform as implemented. Inside the DU process: Class A modules are invoked on the PUSCH lane’s own CUDA stream and write DU-owned outputs; a Class B module is a direct call at the scheduler’s decision boundary returning validated intents; native Class C modules take leased views of each captured slot and publish results into host-owned caches the scheduler reads back. Outside: supervised Class C workers behind a shared-memory ring or CUDA-IPC pool, and E3 clients of every kind. One embedded E3 agent manages all of it off the real-time path. TABLE I T HE THREE EXECUTION CLASSES AT A GLANCE . I NTERFACE IDENTIFIERS ARE THE 32- BIT CONTRACT IDS A PACKAGE DECLARES . Class

Interfaces

Runs

Input

Output

Budget

When it fails or is late

Class A, resident inline L1

0x00010003 channel estimation; 0x00010004 estimation plus equalization; 0x00010005 full receiver to soft bits

in the DU process, enqueued on the PUSCH lane’s own CUDA stream (CPU variants call the CPU upper PHY directly)

GPU-resident slot grid (complex BF16), DM-RS pilot tensor, compact data-RE indices, typed PUSCH and DM-RS metadata

channel estimates plus noise; equalized symbols plus post-equalization noise; FP16 soft bits before descrambling

100 µs for estimation, 150 µs to completion for the deeper two

out-of-profile grant, bypass, or a failed enqueue: the conventional stage runs in the same invocation. Late completion: the result is used, an incident is recorded, and eight in a row open the lane breaker

Class B, bounded real-time control

0x00010006 scheduler

in the DU process, a direct call at the scheduler’s per-slice decision, after the conventional policy has run

up to 8,192 pointer-free candidate records, the latest interference map, the free-VRB mask, and the flags the operator allows

intents per candidate: priority or forbid, PRB limits, preference, or assignment, MCS, layers, TPC, and a per-PRB avoid mask

100 µs from feature construction through commit

too late before the call: skipped. Late or invalid after it: discarded. Late after commit: rolled back. The conventional decision always stands

Class C, asynchronous observation and advisory

0x00010001 full uplink spectrum grid; 0x00010002 SRS channel estimates

in process on a leased view, in a supervised process on the same host, or as an E3 client on any host

one immutable publication per slot: grid or SRS tensor, the scheduler’s expected-use plan, sequence and radio time

per-PRB interference map and quiet-period intents from the one admitted authority package; observations from everyone else

none on the producer; the consumer paces itself

a full ring or a slow consumer is a counted drop; a dead worker is restarted under a bounded budget; the map it published expires by its validity horizon

sequence of integer bounds. The module then enqueues kernels and returns without allocating or synchronizing. Two clocks govern what happens next, and the distinction matters for operators. The callback clock covers validation and enqueue;

a bypass, an unsupported shape, or a failed enqueue on that clock selects the conventional stage, which is ordered on the same stream before the module’s writes would land, so fallback is deterministic without host synchronization. The

TABLE II VOCABULARY USED IN THIS PAPER .

TABLE III T HE C LASS B CONTRACT ( U S E _ C A S E S / V 1/ S C H E D U L E R . H ).

Term

Meaning

Element

Content

lane, worker

a host thread that invokes a dApp; lane state is lock-free and local, instance state is shared across lanes the single atomic that decides whether a hook is invoked; closed means the conventional path runs a module declining an invocation; the host runs the conventional stage the shape bounds a Class A interface declares (maximum PRBs, ports, layers, modulations, DM-RS types); a grant outside it takes the conventional path a Class B module’s advice to the scheduler; validated field by field, committed or discarded atomically an operator YAML switch granting a Class B instance authority over one intent class; all default off YAML scheduler_evaluation_only: a Class B instance is invoked and validated on every decision, its output discarded and counted a lifecycle state: warm, configurable, model-loadable, zero traffic exposure; not invoked a monotonic counter on inventory, each instance, each configuration, and each model; mutations carry the expected value and a stale one returns conflict the bounded lifetime of a Class C view or a data-store subscription; abandonment lets it lapse a per-stream minimum interval no subscriber may go below; non-zero by default a per-instance latch that opens after eight consecutive faults or misses and routes traffic conventionally until an operator probe an attributed record of a fault, miss, breaker transition, or lifecycle change, in a bounded store a disposable helper process that queries a module’s descriptor before the DU maps it proof that no worker is inside a module; required before unload the hash-bound JSON record that is signed, and the monotonic release number folded into the signature the operator’s offline file of vendor keys, per-package sequence floors, revocations, and a policy epoch a host-granted, budget-clamped uplink quiet window requested by a sensing authority the SCTP carriers: APER-encoded management and telemetry on port 36423; a FlatBuffers poll-only data profile on port 38472

Invocation

direction, slice, radio time, cell PRB count, remaining slice budget, candidate generation, allowed output flags candidate id, RNTI, pending bytes and pending SR, head-of-line delay (downlink), served-rate average, the conventional priority and forbid, CQI or SINR, first-transmission BLER over a 64-outcome window with its count, normalized power headroom, the host’s OLLA offset, last outcome and the MCS it used, MCS table and expert limits, layers, transform precoding, target SINR; validity flags per field the latest Class C interference map as a per-PRB tensor with its provenance header; the free-VRB mask priority or forbid; minimum and maximum PRBs; preferred start and count (a soft preference, or a hard assignment when allowed); recommended and maximum MCS; layers; TPC command; beam id (placeholder, no v1 authority) the intent array, a decision generation that must match the input, and an optional per-PRB avoid mask for this slot’s new uplink transmissions avoid_prb_mask (needs a nonzero PRB budget), mcs (bounded by a maximum delta from the host), prb_limits, prb_preference, prb_assignment (needs the free-VRB snapshot), tpc, layers every candidate covered exactly once; every flagged field in range and permitted; storage identities unchanged; all or nothing

gate bypass admission profile

intent allow flag evaluation mode

shadow generation

lease hard floor breaker incident preflight quiescence manifest, package sequence trust policy quiet reservation E3AP, E3DP

completion clock is measured by per-lane CUDA events bracketing the module and its continuation to the pre-descrambling boundary. A late completion is not replayed: the host has already committed to consuming the module’s result for that grant, so the result is used, an incident with the observed latency is recorded, a counter increments, and eight consecutive misses open the lane’s breaker so later grants take the conventional path. Host-side waiting is bounded by a soft wall budget on the order of 2 ms, so a wedged kernel costs at most the current slot. The production budgets are 100 µs for estimation and 150 µs to completion for the deeper two depths. Model management is first class: a bounded stage, validate, warm, activate, rollback sequence swaps weights on a live cell under generation checks with no allocation on the hot path. The reference receiver owns exactly two preallocated device weight banks, publishes the active one with a single atomic exchange, keeps the displaced bank as the only rollback generation, and reuses a bank only after per-lane completion fences prove it drained. B. Class B: bounded real-time control The MAC scheduler’s per-slice policy computes its conventional priorities first and keeps its normal history. It then

Candidate record (one per UE, at most 8,192)

Context Intent (per candidate)

Output

Allow flags (YAML)

Validation

builds a pointer-free candidate snapshot and calls the admitted Class B module directly on its own lane. Table III lists the complete contract as the header defines it: what each candidate record carries, what an intent may say, and which flags grant which authority. The complete boundary, from feature construction through validation and commit, runs under an admitted deadline of 100 µs, a host default that is deliberately not a YAML knob. Overruns are handled at three checkpoints: too late before the call, the call is skipped; late or invalid after it, the result is discarded; late after commit, everything applied is rolled back. In every case the conventional decision stands, and repeated misses open the instance’s breaker. Authority is granted flag by flag. An unset flag makes that intent class silently unavailable, and the module is told so through allowed_output_flags. The reference scheduler is deliberately a baseline replica whose five stages (ranking, outer-loop link adaptation, MCS selection, PRB packing, uplink power control) all ship off, so activating it with every flag set changes nothing until a stage is switched on over E3. The recommended path is evaluation mode first: the module is invoked and validated on every decision, its output is discarded and counted, and its phase timers accumulate under real load. Only then does the operator flip individual flags. C. Class C: asynchronous observation and advisory Class C is where prior architectures ended and where OCUDU deliberately begins. A PHY producer creates one immutable publication per slot: the full uplink grid or the SRS channel estimates, the scheduler’s committed expected-use plan (a symbol-by-PRB category map, so a radiometer can tell planned-empty cells from our own UEs’ energy), a sequence number, and radio time. Three isolation levels consume it: native in-process modules over leased zero-copy views or route-owned snapshots; supervised worker processes over a

shared-memory ring, or over a bounded CUDA-IPC export pool filled by one device-to-device snapshot; and portable E3 clients on any host. Nothing ever maps live L1 memory into another process. The producer never waits: publication is a trylock plus a non-blocking eventfd, one write and two read syscalls, and a full ring or busy consumer is a counted drop. Host-inline Class C copies the full grid on the uplink PHY thread, roughly 50 to 100 µs per slot for a 273-PRB, four-port grid, so multi-cell deployments use the CUDA route, which costs one device-to-device copy instead. Results flow back, but under a strict authority model. A package’s manifest carries a capability bit that distinguishes the single admitted map authority from any number of observer-only workers. Only the authority may publish perPRB interference maps and quiet-period intents, and only through the in-process host API or the supervised worker’s authenticated control channel, whose peer credentials are pinned at spawn. Each result must match a single-use ledger entry recorded when its input was committed, and the host, not the worker, stamps instance identity, generation, and sequence. A result frame from an observer package is a fatal capability violation. E3 clients cannot publish maps at all; they act back only through management operations. The scheduler’s avoidance mask consumes the map and any granted quiet reservation on a later decision, typically within a few slots, and every map carries a validity horizon after which it is ignored. Fig. 3 places all three contracts on one slot. D. A continuum of performance and trust The classes are points on a performance–trust continuum that the design spans on purpose (Fig. 4). A resident Class A module runs in process on the host’s stream with zero-copy tensors and microsecond authority, and must therefore be the most trusted artifact in the system: signed native code in the DU’s address space. An E3 client at the other end runs in any process, language, or host, needs almost no trust, and pays for that isolation in latency, copies, and rate floors. Between them sit Class B’s bounded in-process boundary, native Class C placement, and the supervised process with its one-copy export and seccomp containment. A vendor or integrator enters at the placement that matches its containment requirements and maturity, typically the isolated end, and moves inward as evaluation-mode evidence, A/B windows, and incident-free operation accumulate. Placement is a Class C choice: Class A and committing Class B are in process by definition, and a package declaring an out-of-process inline hook is rejected at catalog time. Section XII reports what each position costs on the reference host, and the companion interface study [14] measures every position against the mechanisms of the prior frameworks. V. PACKAGES , T RUST, AND THE L IFECYCLE What a package is. A production bundle is four files: the shared object, a JSON manifest, an SPDX 2.3 SBOM, and a detached signature sidecar. The manifest names the package identity, the interfaces, backends, and placements it

declares, the SHA-256 of the artifact, and the SHA-256 of the SBOM. One signature covers all of it: the vendor signs a domain-separated digest of the manifest concatenated with the package sequence number, using ECDSA P-256 with SHA256. Verifying the signature and re-hashing the artifact and SBOM against the manifest therefore authenticates the whole bundle, and folding the sequence in gives rollback protection. The SBOM is checked, not decorative: it must describe exactly one package and one file whose name, version, supplier, and digest equal the manifest’s. What the operator holds. An offline trust policy names every trusted vendor public key by the hash of its SPKI encoding, a per-package sequence floor, a revocation list, and a monotonic policy epoch that the gNB YAML pins with a minimum. Rotating a key adds the new key and bumps the epoch; revoking one invalidates every signature it ever made; blocking a bad build raises that package’s floor. The policy file itself is protected by filesystem permissions and the YAML epoch, not by a signature, which is stated plainly in the security model. Admission then binds the artifact hash and SBOM identity, runs descriptor discovery in a disposable preflight process, checks the manifest against the compiled descriptor, and rejects the package before the DU maps it if any of that fails. The lifecycle. Fig. 5 draws the instance state machine as the runtime implements it. Three states are easy to confuse. LOADED means preflighted and mapped, not callable. ARMED means every hook is bound behind a disabled gate and all cold work, including warm-up at the declared worst shape, is done. SHADOW means the instance is warm, accepts configuration and model operations, and sees zero traffic. Shadow is therefore where weights are staged, validated, and warmed with no exposure, and where an A/B candidate waits; it is not where behaviour is measured. Measurement without authority is evaluation mode for Class B, which is ACTIVE with the output discarded, and for Class A it is activation itself, compared across windows. Activation resolves the declared interface set against the hook registry, stages every lane behind one instance-unique gate, and exposes the whole set with a single release store. Deactivation reverses the order: the gate closes, lanes unpublish after a bounded reader grace period, and only then is the module asked to stop and quiesce. A module that cannot prove quiescence is quarantined resident rather than freed under a live kernel. Concurrency and models. Inventory, each instance, each configuration, and each model carry a monotonic generation. Every mutating request carries a transaction id, the expected generation, and an optional dry-run flag; a stale generation returns conflict instead of overwriting, and long operations return a job id instead of holding a socket. Model operations follow a fixed ladder, stage, validate, warm, activate, rollback, retire, over a three-bank store (staged, active, previous) that the SDK implements for every class, so a vendor supplies only the meaning of validate and warm for its own artifact.

14 OFDM symbols of uplink slot n arrive (500 µs at 30 kHz SCS)

try-lock, 1 device copy, 1 eventfd write; a full ring snapshot + publish is a counted drop, never a wait 100 µs admitted deadline

Class C

Class B

Class A

conventional priorities (always computed first)

late or invalid: discarded, conventional result stands

a later decision consumes the map or the quiet grant

CE / EQ / RX kernels on the host’s CUDA stream

budget miss: incident + circuit breaker

invoke, validate, commit

enqueue and return in microseconds; completion clocked by a CUDA event

consumer analyzes at its own pace: native module, supervised worker, or E3 client provenance-checked result

conventional receiver armed as per-invocation fallback slots n+1, n+2, . . .

slot n ends

Fig. 3. One uplink slot across the three timing contracts. Class B decides early inside a bounded boundary; Class A replaces receive stages inline as the slot completes, measured on the completion clock; Class C forks a snapshot without blocking and closes the loop at a later decision through provenance-checked results. performance · data-path efficiency · timing authority

isolation · containment · language freedom · ease of onboarding

Class A resident in-process, host’s CUDA stream zero-copy tensors, µs

Class B in-process bounded 100 µs boundary, validated intents only

Class C native in-process, leased zero-copy views, breaker-contained

Class C process supervised: one-copy shm / CUDA-IPC · seccomp · heartbeats · container-able

E3 external any host, language, container · SCTP · normative schemas, ms

signed native code in the DU: most trust

trusted code, authority allow-flagged

trusted code, fault-contained

crash / address-space isolation, sub-ms

least trust needed; latency + copies traded away

portable, frozen C ABIs — adoptable by other RAN stacks

stack-agnostic today: anything speaking the E3 schemas

typical adoption path: enter at the isolated end, earn trust with evidence (shadow, A/B, incidents), move inward as maturity grows — each SI and vendor picks the point that matches their priorities; both ends are first-class

Fig. 4. The performance–trust continuum. Five placement points span resident in-process execution (maximum authority, maximum required trust) to external E3 processes (maximum isolation and language freedom). The in-process contracts are portable C ABIs; the E3 end is stack-agnostic.

VI. T HE E3 P LANE Everything is controlled and observed over E3 (Fig. 6) through published, versioned schemas: an ASN.1 management service model and an E3AP envelope encoded as APER, a telemetry service model, and a FlatBuffers data profile. Three carriers serve them. A credentialed local UNIX socket, mode 0600 with peer credentials checked, serves same-host tooling. E3AP over SCTP on port 36423 carries setup, subscriptions, indications, and management control with message-boundary framing and bounded fragmentation up to 16 MiB. The data profile (E3DP) over SCTP on port 38472 is a FlatBuffers polland-lease endpoint for consumers that do not want an APER stack. The APER decode path is defensively bounded end to end (constraint checks before field reads, bounded reassembly, size-capped strings), so a malformed peer cannot destabilize the agent.

E3 here is a real-time local evolution of E2’s setup, subscription, indication, and control philosophy, bridgeable to E2 but not a ratified O-RAN interface. Table IV states exactly how it relates to the two prior implementations. The procedure names and field vocabulary were kept from libe3 so that an adapter is mechanical, but the wire formats are not bytecompatible, and a libe3 dApp needs an envelope re-encode plus a channel-topology shim rather than a direct connection. SCTP is the carrier E2 standardized and the one several large vendors prefer for RAN control planes, which is why it is a first-class carrier here. Transport security, stated plainly: E3 carries no TLS, DTLS, or token authentication in this release. Both SCTP carriers bind to loopback by default and admit every peer that reaches the bound address unless a literal peer list is configured. A management association carries the full authority

DISCOVERED

load

LOADED

prepare

in the catalog, preflighted, verified, not mapped dlopen’d; not callable

PREPARED

arm

ARMED

enterShadow SHADOW

activate

all allocation done; hooks bound, gate zero traffic; config weights uploaded disabled; warm-up done and model ops legal

ACTIVE

deactivate

gate open; the radio sees the dApp

UNLOADED

unload

DRAINING

gate closed, lanes detached

INACTIVE

prepare again: re-arm the same instance degraded: a lane failed to bind or a worker died; the conventional path is in use

circuit breaker (per instance, across lanes) is not a lifecycle state: eight consecutive faults or misses open it, traffic routes conventionally, health shows it, breakerProbe closes it. Evaluation mode (Class B, YAML) is ACTIVE with the output discarded.

quarantined-resident: quiescence could not be proven; image held, never freed under a live kernel; retryQuiescence

every transition is an asynchronous job carrying a transaction id and the expected instance generation; a stale generation returns conflict; every action accepts dryRun

Fig. 5. The instance lifecycle as implemented. E3 actions (violet) drive the transitions; each state’s meaning is written under it. The circuit breaker and Class B evaluation mode are orthogonal to the state and are shown as notes. Every transition is an asynchronous, generation-checked, dry-runnable job.

producers

embedded E3 agent

carriers

clients

8 telemetry capture sites (PHY threads) pusch.symbols, pusch.crc, pusch.grid, pusch.dmrs_channel, srs.channel, gnb.metrics_json

portable data store bounded slots and byte budget; leases; try_lock + counted drops

local UNIX socket SO_PEERCRED checked

dappctl (C++) ocudu_dappctl.py --sctp

Class C snapshot routes classc.spectrum, classc.srs_isac

management service model inventory, lifecycle jobs, typed config, model slots, metrics, incidents, health, quiet reservations, cell info, RAN control

before any capture work: idle = 1 atomic load; per-RNTI and streamwide floors consumed by CAS

3 monotonic generations (inventory, instance, configuration); dryRun on every mutation; transaction ids for idempotent retry

E3AP over SCTP 36423 ASN.1 APER; bounded decode; 64 KiB fragments E3 data profile, SCTP 38472 FlatBuffers poll + lease

python3 -m ocudu_e3 MCP server, 28 tools (LLM agents) recorders, replay, dashboards, dApps

control returns under host authority: lifecycle, configuration, model swap, quiet-period grants, UE release and handover, RU gain

Fig. 6. The E3 plane. Producers pay one atomic load per event when nobody listens and consume rate floors by CAS before any capture work. One embedded agent with a bounded store serves three carriers; control returns under host authority.

TABLE IV W IRE COMPATIBILITY WITH THE PRIOR E3 IMPLEMENTATIONS . OCUDU E3

libe3 (NEU)

NVIDIA cuSense

Encoding

ASN.1 APER

JSON

Transport

SCTP with message boundaries

Envelope

E3-PDU, same procedure names 16 MiB with fragmentation response with a service-model body the published schemas

ASN.1 APER or JSON ZMQ or POSIX sockets, three channels E3-PDU

Payload bound Control outcome What an existing dApp needs

32 KiB, no fragmentation acknowledgement only an envelope adapter; E3DP serves the reference client’s transport, but its shipped schema must be regenerated

ZMQ REQ/REP and PUB/SUB JSON common_headers out-of-band shared memory JSON response a new transport layer

of the service model, including load, unload, configuration, and RAN control, so an operator must treat the E3AP port exactly like the local management socket and place an external boundary (IPsec, VPN, network policy) in front of any off-host exposure. Telemetry (Table V) differs from prior exports in two

properties. Rate discipline: per-RNTI floors, per-subscription thinning, and a non-zero-by-default per-stream hard floor are enforced before any capture work, and an idle stream costs one atomic load per event, so subscribing a recorder cannot destabilize the cell; full-rate capture is an explicit operator optin. Completeness: pusch.grid carries raw pre-equalization resource elements with every parameter needed to re-run the receive chain (HARQ and RV, scrambling, LDPC base graph, MCS, DM-RS layout, DC position, UCI), the substrate for replay-verified datasets. Management is one surface. It reports inventory, metrics and metric schemas, incidents, health, served-cell information, and quiet-reservation queries. It drives asynchronous lifecycle jobs, typed configuration against per-package schemas, model operations, and a bounded set of RAN controls (UE release to idle, handover, RU gain) that the host’s own authority object admits and defers to the owning component. The same operations work identically over the local socket and SCTP, from the C++ dappctl, the Python client, or the MCP server of Section VIII. VII. W HERE THE C ODE L IVES The platform is published as three repositories under the OCUDU AI-RAN Working Group 2 (Fig. 7), cloned side

ocudu-dapp-platform (clone as installed headers only scaffold, certify ocudu/) the gNB with the embedded dApp runtime ocudu-dapp-sdk and E3 agent the authoring kit; builds against the a vendor repository (yours) include/ocudu/dapp/ the frozen ABI installed platform package only (abi/v1), the 6 use-case interfaces scaffolded by the SDK; reference/ 16 reference packages (use_cases/v1), the process contracts find_package(OCUDUDAppSDK); (CE, EQ, receiver, scheduler, spectrum, (process/v1) own CI, own cadence, own models SRS-ISAC; CPU and CUDA), supervised lib/dapp/ runtime, E3 agent and output: shared object + signed manifest workers, portable E3 consumers codecs, IPC + SPDX SBOM examples/ vendor_min, Python Class lib/phy/upper/dapp/, C over E3, telemetry observers, sensing, admitted by the operator’s trust policy, lib/scheduler/dapp/ the Class A and MCP server named in the gNB YAML, driven over E3 B seams tools/ certify, scaffold, validate; docs/ apps/ dApp service, dappctl, packagauthoring guides ing, preflight schemas/e3/, python/ocudu_e3/, docs/dapp/ ocudu-dapp-quickstart (clone as ocudu-docker/): 1 Dockerfile builds both repositories and runs every release gate; CPU and CUDA images with radio support (UHD, DPDK); docker-compose.yml with a loopback E3 agent so every tutorial runs without hardware; the tutorial ladder and documentation map

the seam everything builds against: installed CMake package OCUDUDAppSDK, frozen C ABI v1 with a checked-in layout fingerprint, 6 versioned use-case interfaces, published E3 schemas. No repository links against another; packages meet only at runtime, through the host.

Fig. 7. The three public repositories and the seam between them. The platform repository owns the interfaces, the runtime, and the E3 plane; the SDK and every vendor repository build against the installed package alone; the quickstart builds all of it into a container with a loopback E3 agent.

TABLE V T ELEMETRY STREAMS AND THEIR DEFAULT HARD FLOORS . Stream

Content

Floor

pusch.symbols

Equalized symbols, SINR, EVM, EPRE, RSRP, TA, CFO, grant shape Per-transport-block CRC verdict Per-subcarrier LS SRS response per rx/tx port pair DM-RS channel estimate per port and layer Raw pre-EQ REs plus all parameters for offline re-decode The gNB JSON metrics report, byte for byte Class C snapshot: full uplink grid tensor and expected-use plan Class C snapshot: SRS channel-estimate tensor

10 ms

pusch.crc srs.channel pusch.dmrs_channel pusch.grid gnb.metrics_json classc.spectrum classc.srs_isac

1 ms 5 ms 5 ms 20 ms report period producer rate producer rate

by side with the platform as ocudu/ and the quickstart as ocudu-docker/. The working group develops them in the open. This release is a preview, so others can test the interfaces, propose use cases, and vet the references before the runtime is upstreamed into the OCUDU mainline. Table VI maps each mechanism in this paper to its entry point, so a reader can go from a sentence here to the code that implements it. ocudu-dapp-platform is the gNB with the embedded runtime. The public ABI is include/ocudu/dapp/ abi/v1/ (module, host API, memory, execution, management, metric catalog), the six use-case interfaces are include/ocudu/dapp/use_cases/v1/ together with coordination.h for the quiet-reservation intent, and the out-of-process contracts (shared ring layout, CUDA export pool, SEQPACKET control protocol) reside in include/ ocudu/dapp/process/v1/. The runtime under lib/ dapp/runtime/ holds the package catalog and manifest,

the artifact verifier and signature checks, the disposable preflight, the module loader, the native hook registry and slots, the instance manager and lifecycle, the job dispatcher, the circuit breaker, the quiescence tracker, and the Class C process supervisor. The E3 plane under lib/dapp/e3/ holds the embedded agent, the management service, the APER and FlatBuffers codecs, the E3AP and E3DP SCTP servers, the local credentialed transport, the portable data store, and the telemetry publisher; lib/dapp/ipc/ holds the shared SPSC ring, the SEQPACKET channel, and the Class C publication routes. The seams into the PHY and MAC are lib/phy/ upper/dapp/ and lib/scheduler/dapp/. The gNBside service and YAML configuration are apps/services/ dapp/. The operator tools dappctl, dapp_package, and dapp_preflight are under apps/tools/. The schemas are schemas/e3/, the Python client is python/ocudu_ e3/, and docs/dapp/ holds the thirteen tutorials, the glossary, and the reference documents. ocudu-dapp-sdk is the authoring kit. It consumes the installed OCUDUDAppSDK CMake package and never copies host headers. reference/ holds sixteen reference packages (channel estimator, equalizer, receiver, scheduler, spectrum, SRS-ISAC, in CPU and CUDA variants), the supervised process/ workers, and the portable/ E3 consumers; examples/ holds the copy-me commercial Class A layout vendor_min, a Python Class C dApp over E3, Python telemetry observers, spectrum and ISAC sensing over E3AP, and the MCP server. tools/ holds the certifier, the scaffolds for every class, and the validation ladder. docs/ holds the per-class authoring guides. ocudu-dapp-quickstart is a Dockerfile that builds both repositories from their current commits, runs every release gate inside the build, and produces CPU and CUDA images with radio support (UHD, DPDK) baked in, plus a

TABLE VI W HERE EACH MECHANISM LIVES ( PATHS RELATIVE TO THE PLATFORM REPOSITORY UNLESS MARKED SDK).

TABLE VII T HE TUTORIAL LADDER ( D O C S / D A P P / IN THE PLATFORM , D O C S / IN THE SDK).

Mechanism

Entry point

#

Tutorial

Covers

Frozen C ABI v1

include/ocudu/dapp/abi/v1/module.h, host_api.h, memory.h, execution.h; fingerprint gate in the SDK CI include/ocudu/dapp/use_cases/v1/ (spectrum.h, srs_isac.h, channel_estimator.h, equalizer.h, receiver.h, scheduler.h, coordination.h, pusch_metadata.h) lib/phy/upper/dapp/; hook slots in lib/dapp/runtime/native_hook_*.cpp lib/scheduler/dapp/; intent validation and allow flags in the scheduler policy lib/dapp/ipc/publication_route.cpp, shared_spsc_ring.cpp, seqpacket_channel.cpp; lib/dapp/runtime/ class_c_process_supervisor.cpp lib/dapp/runtime/package_catalog.cpp, package_manifest.cpp, package_signature.cpp, artifact_verifier.cpp, native_preflight.cpp; docs/dapp/security_model.md lib/dapp/runtime/instance_manager*.cpp, lifecycle.cpp, job_dispatcher.cpp, quiescence_tracker.cpp circuit_breaker.cpp, guarded_call.h, incident_store.cpp lib/dapp/e3/embedded_agent.cpp, management_service.cpp, aper_codec.cpp, e3ap_codec.cpp, fbs_codec.cpp lib/dapp/e3/local_connector.cpp, e3ap_sctp_server.cpp, e3dp_sctp_server.cpp; docs/dapp/e3_wire.md lib/dapp/e3/telemetry_publisher.cpp, portable_data_store.cpp schemas/e3/ (ASN.1 management, telemetry, E3AP; FlatBuffers data profile) apps/services/dapp/; reference in docs/dapp/yaml_configuration.md apps/tools/dappctl, dapp_package, dapp_preflight; python/ocudu_e3/ docs/dapp/glossary.md, docs/dapp/known_limitations.md reference/: channel_estimator, equalizer, receiver, scheduler, spectrum, srs_isac (each with a _cuda variant), process, portable tools/: certify_class_a_cuda.cu, scaffold_class_a_cuda.py, scaffold_class_bc_cpu.py, dapp_validate.py examples/e3_mcp_server/; guide in docs/mcp_companion.md

0

Zero-hardware testbed

1

New vendor dApp repo (SDK)

2

Install and load from the gNB YAML

3

Operate with dappctl

4

Monitor and tune over MCP (SDK)

5

Signing, SBOMs, keys

6

Exposing dApp KPIs over E3

7

First dApp: standalone Class C

8

Telemetry and data capture

9

Neural-model operations

10

Class B scheduler control

11

Troubleshooting runbook

12

Capstone

build the images; loopback E3 agent; run everything radio-free scaffold, build, certify against the installed SDK catalog, trust policy, initial instances, admission lifecycle, A/B version swap, deactivate and unload, breakers ask the gNB questions; KPI-driven LLM tuning loops what is signed, the SBOM, trust policy, rotation, revocation declare metrics; read them from any client a Python out-of-process dApp: subscribe, decode, act back all eight streams, floors, HDF5 capture, offline PUSCH replay stage, validate, warm, activate, rollback; live A/B the 100 µs contract, allow flags, evaluation-first, sensing to avoidance symptom, cause, fix, from real failure modes capture, train, package, shadow, activate, monitor, rollback

Use-case interfaces

Class A seam Class B seam Class C native and process

Admission and trust

Lifecycle and jobs

Safety E3 agent and codecs

Carriers

Telemetry Schemas YAML and service Tools Glossary, limitations References (SDK)

Certify and scaffold (SDK)

MCP server (SDK)

Compose file with a loopback E3 agent so every tutorial command runs without hardware. VIII. W ORKING W ITH THE P LATFORM The first session. The zero-hardware path builds the platform into a container and points every later command at a loopback E3 agent. A green image is a validated platform, because every release gate runs inside the build. # 1. Clone the three public repositories side by side WG2=https://gitlab.com/ocudu/work_groups/wg2_ai_ran git clone $WG2/ocudu-dapp-platform.git ocudu git clone $WG2/ocudu-dapp-sdk.git git clone $WG2/ocudu-dapp-quickstart.git ocudu-docker # 2. Build the release images (host, SDK, every gate) ocudu-docker/scripts/build.sh # 3. Start the loopback E3 agent and look around

docker compose -f ocudu-docker/docker-compose.yml up e3agent python3 -m ocudu_e3 --port 36423 catalog python3 -m ocudu_e3 --port 36423 subscribe pusch.crc -maximum 5 python3 -m ocudu_e3 --port 36423 inventory

The tutorial ladder. The documentation is organized as thirteen tutorials (Table VII) that take a reader from an empty vendor repository to an LLM agent tuning a live cell. The first five need no signing keys, and all of them run against the loopback agent. What a module looks like. Every in-process dApp, whatever its class, exports exactly one symbol, ocudu_dapp_query_v1, which fills a module table: identity, the interfaces it implements, and the non-real-time lifecycle callbacks create, configure, prepare, warmup, start, stop, quiesce, destroy. Each interface supplies one real-time entry point. Listing 1 shows a complete, if minimal, Class B module condensed from the SDK’s reference scheduler: it echoes the conventional priority for every candidate, which is what makes it a baseline replica, and lowers the MCS of a UE whose recent block error rate is high, but only when the operator has granted MCS authority. Every structure begins with its size and ABI version, no pointer to a host object crosses the boundary, and the module reads only the size it was given. The SDK’s helper layer fills the module table and the worker context from a few lines of C++, so a vendor writes the algorithm, not the plumbing. #include "ocudu/dapp/use_cases/v1/scheduler.h" #define SCH(x) OCUDU_DAPP_SCHEDULER_##x /* shorten the ABI names */ static ocudu_dapp_status_v1 invoke(void* ctx, const ocudu_dapp_scheduler_input_v1* in, ocudu_dapp_scheduler_output_v1* out)

{ /* in->invocation.deadline_monotonic_ns bounds this call */ uint32_t n = 0; for (uint32_t i = 0; i < in->nof_candidates; ++i) { const ocudu_dapp_scheduler_candidate_v1* c = &in-> candidates[i]; ocudu_dapp_scheduler_intent_v1* it = &out->intents[n++]; it->struct_size = sizeof(*it); it->candidate_id = c->candidate_id; it->flags = SCH(INTENT_PRIORITY_V1); it->priority = c->conventional_priority; /* baseline replica */ if ((in->allowed_output_flags & SCH(ALLOW_MCS_V1)) && (c->flags & SCH(CANDIDATE_BLER_VALID_V1)) && c->bler > 0.2f && c->recommended_mcs > 0) { it->flags |= SCH(INTENT_MCS_V1); it->recommended_mcs = (int16_t)(c->recommended_mcs 1); } } out->nof_intents = n; out->decision_generation = in->candidate_generation; return OCUDU_DAPP_OK_V1; /* host validates every field, then commits */ } /* interface table: worker-context hooks + one real-time entry point */ static const ocudu_dapp_scheduler_interface_v1 api = { WORKER_API, invoke, {0} }; /* the only exported symbol; the descriptor names interface 0x00010006, Class B, in-process placement, per-worker threading */ OCUDU_DAPP_EXPORT_V1 ocudu_dapp_status_v1 ocudu_dapp_query_v1(const ocudu_dapp_host_info_v1* host, ocudu_dapp_module_v1* m) { return populate(host, m, "com.vendor.sched", &descriptor, &api); }

Listing 1. A minimal use_cases/v1/scheduler.h.

Class

B

module

against

Building and certifying a dApp. A vendor scaffolds a project with scaffold_class_a_cuda.py or scaffold_class_bc_cpu.py, builds it with find_package(OCUDUDAppSDK) against the installed package, and runs certify_class_a_cuda (for Class A) or dapp_validate.py. The certifier loads the final shared object through the public ABI, runs the complete lifecycle, creates two workers on independent CUDA streams, and drives deterministic vectors from the one-PRB minimum to the package’s declared maximum. Two of those vectors matter especially: off-origin placements catch a module that interprets DM-RS coordinates relative to its own allocation, and rank-deficient two-layer references resolve only when the module consumes the joint multi-symbol pilot set. Every caller-owned output sits at a nonzero offset inside a poisoned allocation with prefix, suffix, and inactive-capacity canaries checked after every call. dapp_package then emits the manifest and SBOM and signs the bundle with the vendor’s key. Loading and driving a dApp. The operator stages the signed bundle in the catalog directory named in the gNB YAML and drives it over E3. The YAML names manifests only; E3 can select from the catalog but never supply a path. # gnb.yml: the dapp block (production shape) dapp: catalog_root: /opt/ocudu/dapps

native_preflight_helper: /opt/ocudu/bin/ocudu-dapppreflight package_signature_mode: required # lab: allow_unsigned package_trust_store: /etc/ocudu/dapp-trust.json minimum_trust_policy_epoch: 1 package_manifests: [acme_eq.manifest.json, acme_sched. manifest.json] initial_instances: - {package_manifest: acme_eq.manifest.json, placement: in_process, backend: cuda, target: shadow} # warmed, zero traffic scheduler_evaluation_only: true # measure before authority scheduler_allow_mcs: true # one allow flag per intent class local_e3: {socket_path: /run/ocudu/dapp-management.sock} e3ap_sctp: {bind_address: 127.0.0.1, port: 36423} # walk the lifecycle over the local socket; every step is an async job ocudu-dappctl inventory ocudu-dappctl load --package <id> --placement in-process -backend cuda ocudu-dappctl lifecycle prepare <instance> # allocate , upload weights ocudu-dappctl lifecycle arm <instance> # bind hooks, gate closed ocudu-dappctl lifecycle enter-shadow <instance> # warm, zero traffic ocudu-dappctl model <instance> stage|validate|warm|activate |rollback ocudu-dappctl config <instance> --set olla_target_bler=f64 :0.05 --dry-run ocudu-dappctl lifecycle activate <instance> # one release store ocudu-dappctl metrics <instance>; ocudu-dappctl incidents; ocudu-dappctl health # the same operations over SCTP from any host python3 ocudu_dappctl.py --sctp <gnb>:36423 inventory

Two instances of competing packages can be loaded, one active and one waiting in shadow with its weights already warm, and swapped by activating one and deactivating the other. The KPI comparison then runs across activation windows, which is how equalizer variants were compared over the air (Section XII). Writing a portable Class C dApp. A Python process on any host opens an E3AP association, subscribes to classc.spectrum or srs.channel, decodes each indication into a NumPy array, and acts back through management operations. Listing 2 is the core of the SDK’s emitter detector, which was active in the composition run of Section XII. On a 51-PRB cell it keeps up with the full producer rate, about 120 grids per second at 4 MB/s, and decodes each grid in about 120 µs. Nothing in the gNB is patched or rebuilt. from ocudu_e3 import E3Client, SubscriptionConfig from ocudu_e3.classc import decode_spectrum with E3Client("gnb-host", 36423, dapp_name="radiometer") as client: client.setup() # E3AP handshake; assigns a dApp id cfg = SubscriptionConfig("classc.spectrum", drop_policy= "dropOldest") client.subscribe(["classc.spectrum"], configs=[cfg]) for ind in client.indications(timeout=1.0): rec = ind.record # OCC3 header + raw grid bytes snap = decode_spectrum(rec.metadata, rec.tensor) p_db = 10 * np.log10(snap.prb_power() + 1e-20) # [ symbol, PRB] # track a floor only where the plan says "planned empty";

# report PRB runs above it that persist across slots for event in detector.detect(snap): print(event, ind.dropped_objects)

Listing 2. A portable Class C (examples/e3_spectrum_sensing).

radiometer

over

E3AP

Driving the cell from an agent. The MCP server in the SDK (examples/e3_mcp_server/) is an E3 client that exposes the interface’s KPIs and knobs to LLM agents in the simplest possible form: 28 Model Context Protocol tools over one E3 association, twenty that observe (sector state, per-UE KPIs, traffic history, SRS channel, constellation, telemetry catalog, dApp inventory, health, metrics, configuration, incidents, jobs, quiet reservations) and eight that act (dApp configuration and model operations, lifecycle, UE release and handover, RF gain). Every state-changing tool dry-runs by default; applying a change is a second, explicit call with dry_run=false, and a --read-only flag removes the acting tools entirely. Listing 3 condenses a session recorded against the reference cell on 26 August 2026. The agent needed no protocol knowledge, and the platform’s guardrails (typed validation, generations, allow flags, host-arbitrated RAN control) bounded every call: out-of-range values were refused by the dApp’s own validator and a stale generation by the host, both exercised live. user > How is the sector looking? agent> get_cell_overview() -> 4 RNTIs in 10 s; UL SINR -10..9.7 dB; UL BLER mean 0.52, worst 1.0 agent> get_kpis(45); list_ues() -> 0x4711: SNR 22 dB, CQI 15, UL MCS 15-18, BLER 10-12% on 4000 TBs -> 0x474a, 0x474f, ...: SNR -10 dB, MCS 0, BLER ˜100%, a few hundred TBs agent: one real user sitting on the host OLLA’s 10% target; the rest are a device failing attach at the cell edge. An RF problem, not scheduling. user > Can we tighten link adaptation for the good UE? agent> get_dapp_config(instance=sched, schema=true) -> 22 keys, all passthrough agent> set_dapp_config(instance=sched, dry_run=true, {olla_mode: contained_baseline, mcs_mode: table_plus_olla, olla_target_bler: 0.10, olla_step_db: 0.05, olla_max_offset_db: 6.0}) -> dryRunSuccess agent: proposal validated; apply with dry_run=false when you are ready, then compare BLER and MCS over a comparable window and keep or revert.

Listing 3. An LLM agent session over MCP, condensed (reference cell, 26 August 2026).

IX. S EVEN U SE C ASES ON O NE A RCHITECTURE Each use case below is stated as what runs, which path it exercises, and what the platform provides. All of them ran on the live cell of Section XII except the continual-learning loop, whose end-to-end campaigns ran on a dummy-RU DU. Spectrum sensing descends directly from the original dApp demo and runs here at three escalating trust levels. A Python E3 observer subscribes to the Class C spectrum stream, functionally the Northeastern demo but over published schemas, under rate floors, and with no patched gNB. The same radiometer then runs as a supervised process dApp (ref_spectrumd) under heartbeats and seccomp. Finally the worker acts as authority, returning a per-PRB interference

map and quiet-period intents that appear in the scheduler’s avoidance mask within a few slots. One algorithm, three containment levels, no DU change. ISAC rides the SRS channel the gNB already estimates: srs_isac_monitor.py sketches a range–Doppler map from srs.channel; the packaged Class C references compute a bounded channel-impulse-response preview from the same tensors, the CUDA variant on the producer’s stream without ever materializing it on the host. Dataset capture is a first-class product. record_telemetry.py writes any stream mix to HDF5; export_pusch_burst.py captures pusch.grid bursts; ocudu-pusch-replay re-decodes them offline with the gNB’s own receiver and cross-checks the recorded CRC verdict. Every capture carries its own verified label, channel estimates, and exact PHY configuration. The continual-learning loop (Fig. 8) assembles those pieces. Telemetry feeds an online tuning process, which may run as a dApp, a RIC application, or an edge job. That process returns a sealed model artifact over the same E3 surface. The artifact is staged, validated, and warmed on a shadow instance with zero exposure, then activated atomically under a generation check. Per-model metrics and incidents monitor it; degradation triggers a one-operation rollback; drift re-enters at capture. Every step is a typed, transaction-id’d, dry-runnable operation, so the loop can be driven by a script, a CI pipeline, or an agent. The inline neural receiver is the flagship Class A application. The SDK’s CUDA references implement the exact contracts a neural receiver adopts at the three depths. They also carry the model machinery a learned implementation needs: prepare-time weight upload, per-worker lane-local history, and two bounded device weight banks with completion-fence discipline, so an inactive bank is never freed while a stream still reads it. The neural equalizer that ran on the cell is not in the SDK. It is a package built out of source tree, against the public SDK alone, by a separate team, which is the seam this paper claims, demonstrated. The inline learned scheduler. A learned policy replaces stages behind the same intent schema as the baseline replica and inherits the same evaluation-first path, the deadline, the post-equalization noise from Class A, and the interference context from Class C, so it can be interference-aware on day one. Agent-driven operations. Because every operation is a typed request defined by a published schema, exposing the whole surface to an LLM agent takes only a thin adapter. An agent can run a standing loop, read KPIs and incidents every few minutes, reason about trends, and adjust cell parameters or dApp configuration from a single natural-language prompt, with the same guardrails that make the loop safe for a script. X. T HE V ENDOR E COSYSTEM : B UILD , L OAD , C OMPOSE Everything in Section IX was one team’s applications. The same infrastructure supports teams that have never met (Fig. 9) through one clean seam: the host repository owns the six

live cell — model vn active (Class A / B / C dApp)

atomic activate (generation-checked)

continuous monitoring drives the loop: per-model metrics, fallback counters, incidents, KPIs

live telemetry over E3 pusch.grid/symbols · srs.channel rate-floored, per-RNTI, idle = 1 atomic

regression ⇒ rollback (previous model retained)

continual AI-RAN model lifecycle — one E3 surface end to end — every step is a typed, transaction-id’d, dry-runnable management operation

stage · validate · warm on the shadow instance (zero traffic exposure)

sealed weights vn+1 versioned model artifact

training data streaming subscription or recorded corpora; CRC-verified ground truth

online tuning process a dApp, RIC app, or edge job — streaming updates or batch retrain

Fig. 8. The continual-learning loop. Rate-floored telemetry streams into an online tuning process; sealed model artifacts return over E3, staged and warmed on a shadow instance, and activated atomically; monitoring drives rollback or the next round. One E3 surface carries the whole loop. independent vendor repositories Vendor A — neural receiver Class A CUDA, own models find_package(OCUDUDapp)

signed packages nrx.so · manifest-v1 · SPDX SBOM · detached sig

operator / SI

one OCUDU DU, at runtime

offline trust policy per-vendor keys · epochs · revocation · sequence floors

A: Vendor A neural RX admit + lifecycle

Vendor B — scheduler policy Class B intents pipeline

sched.so · manifest · SBOM · sig

Vendor C — sensing / ISAC Class C suite + E3 apps

sense.so · manifest · SBOM · sig

each repo: own CI, own release cadence, scaffold-generated start, offline certification — no OCUDU source needed

immutable, hash-bound artifacts; SBOM identity independently checked at admission

catalog /opt/ocudu/dapps YAML names manifests only

tuning knobs allow flags · budgets · floors · placement · shadow-first A/B

B: Vendor B scheduler

C: Vendor C sensing (×N ) loaded / swapped / unloaded live over E3 — no DU rebuild, no restart

host arbitrates: no cross-vendor linkage; cooperation only via host-owned caches; conventional fallback always armed

Fig. 9. The ecosystem the infrastructure enables. Vendors build signed packages in their own repositories against the installed SDK; operators and integrators hold the trust policy, compose packages from several vendors in one catalog, and load, tune, swap, and unload them on a live DU: no rebuild, no cross-vendor linkage, host arbitration throughout.

interfaces, the E3 plane, and the runtime; the installed SDK package, the frozen ABI, and the published schemas are the only things anyone builds against. Every producer of dApps sits above that seam with the same headers, certification tools, and packaging. Vendors: your repository, your cadence. A dApp lives in the vendor’s own repository, built against the installed SDK with find_package: no host source tree, no fork, no patch queue. The fingerprint-verified ABI and the published compatibility policy decouple the vendor’s release cadence from the host’s. Offline certification and the validation ladder give the vendor a proof to ship beside the artifact, an immutable, hash-bound bundle of shared object, manifest, SPDX SBOM, and detached signature. Operators and integrators: compose, tune, optimize. The operator holds the other half of the trust relationship: the offline trust policy of Section V. Composition is a catalog decision, signed bundles from several vendors named in YAML and driven over E3, and it is safe by construction because

packages cannot link against or call each other, cooperate only through host-owned caches, and every authority is individually granted. The operator therefore tunes the combination: allow flags decide how much scheduling authority a policy gets, budgets and floors decide how much real-time and telemetry cost a sensing suite may impose, placement decides its containment, and shadow-staged candidates with per-instance metrics make A/B between competing packages routine. A vendor-neutral home, and a path to standards. The code lives under the OCUDU AI-RAN Working Group 2 as a preview release, and the working group is the venue in which it is being vetted, extended with new use cases, and prepared for upstreaming into the OCUDU mainline. Interfaces defined in code, exercised by shipping references, and stress-tested by vendors interworking through them let technical consensus form in the project, at the speed of working software. What then goes to the SDOs, O-RAN’s E3 work [11] and ultimately 3GPP, is a proven artifact. The Class A and Class B C ABIs are deliberately portable beyond this codebase, and the E3 end

already admits any stack that speaks the schemas. XI. T HE L ATENCY A RCHITECTURE AND S AFETY A slot is 500 µs at 30 kHz subcarrier spacing, and at five-cell CUDA L1 scale roughly 100 µs of host jitter already produces fronthaul lateness, so the platform must add nothing measurable to the radio’s steady state. The design guarantees this by construction (Fig. 10) in two moves. First, every steadystate mechanism on a real-time thread costs nanoseconds to single-digit microseconds: an atomic gate when idle, a CAS for a rate floor, a lock-free cache read, a try-lock publication with an eventfd write. Second, work is placed by thread, not by module: the real-time lane’s entire E3 duty is one gated memcpy, while encoding, socket I/O, and lifecycle jobs live on management executors that are allowed to be slow. Five rules make this concrete: 1) Idle is an atomic load. Every capture site and dApp hook is gated by a single relaxed or acquire atomic. An unconfigured feature has no measurable cost. 2) Producers never block. Try-locks, non-blocking eventfds, and counted drops replace every wait a consumer could induce. A slow observer is its own problem, never the radio’s. 3) Every residual wait is bounded. A degraded GPU or wedged consumer costs at most the current slot: fallback fires, a counter increments, no thread stalls. 4) No steady-state allocation on any invocation path: preallocated workspaces, lock-free snapshot pools, builders cleared per encode. 5) GPU discipline. Class A enqueues on the host’s stream and returns. Class C CUDA export is one device-todevice snapshot into a bounded IPC pool, never a mapping of live L1 memory. What a deadline means, per class. The deadlines are enforced differently, and an operator should know which is which. For Class B, the deadline is enforced at three checkpoints and a late result is skipped, discarded, or rolled back, so the radio never sees it. For Class A, a failure on the callback clock selects the conventional stage in the same invocation, but a kernel that has been enqueued cannot be cancelled, so a late-but-successful completion is used and recorded rather than replaced; the containment for a slow module is the breaker, which opens after eight consecutive misses, together with the admission gate that keeps out-of-profile grants away from it. Preempting native code mid-invocation cannot be done safely, and this is a deliberate v1 decision: a package whose timing is not yet trusted belongs in Class C process placement, not inline. For Class C, there is no producer-side deadline at all. Fig. 11 shows the ladder a package climbs and the rail everything falls back to. Admission binds the artifact hash and SBOM identity, runs descriptor discovery in a disposable preflight process, and checks declared shape profiles before any dispatch. Runtime guards catch exceptions at the module boundary and translate them into sanitized incident records, meter deadlines on separate callback and completion clocks, and open per-lane circuit breakers on faults, with

TABLE VIII M ECHANISM COSTS AND DEADLINE COMPLIANCE (GB10 HOST ). RUNTIME ROWS ARE THE PLATFORM ’ S OWN CHECKPOINTS ; SUITE ROWS ARE FROM THE COMPANION INTERFACE STUDY [14], QUIET HOST UNLESS STATED . Claim

Result

Conditions

Reproduce with

Class B boundary (runtime)

3.6 µs P99.9, no misses

Class B direct call (suite)

0.29 µs P99.9 quiet; 5.2 µs with the cell on air

worst profile, 9,000-call development matrix; host without approved CPU controls the same 3.2 KB request over SCTP loopback: 16 µs quiet, 181 µs on air, 0.35 % misses at 100 µs

Class A resident receiver

82 µs P50, 112 µs P99.9 (suite); 92.5 and 105.1 µs (runtime)

platform tests, any x86 or Arm host benchmark suite; on-air rows need the radio certifier on a CUDA host

Class C ring publication

2.0 µs P50, 2.3 µs P99.9 per 68 KB slot

E3AP over SCTP

10 µs one way at 1.5 KB; 1.4 ms at 734 KB; 11 µs control round trip

live shape, 51 PRB, two layers, 256-QAM, exclusive GPU; 150 µs budget; 273-PRB kernels not yet qualified producer critical section; full ring is a counted drop; 0.87 Gb/s sustained across forced worker restarts asn1c APER codec, loopback association, two pinned processes

benchmark suite, any host

quickstart loopback agent

the conventional path in place at every rung. Supervised Class C workers add authenticated SOCK_SEQPACKET control with pinned peer credentials, pointer-free results validated at the trust boundary, seccomp deny-lists, resource limits, challenge–response heartbeats, and a backed-off, healthyuptime-windowed restart budget. Observability closes the loop: per-reason fallback counters, deadline misses with measured GPU durations, heartbeat latencies, restart and incident records, all queryable over the same E3 surface a dApp is managed through. In-process modules remain fully trusted code; preflight and guards reduce the impact of predictable faults and are not a sandbox, and process placement is bounded fault containment, not hostile multi-tenant isolation. XII. M EASURED E VIDENCE AND L IMITATIONS Testbed. All live-cell results come from one over-the-air cell: an NVIDIA GB10 (DGX Spark, Arm) host running the CUDA-resident L1, a split-8 USRP B210, band n78 at 3410.1 MHz, 20 MHz (51 PRB) TDD at 30 kHz subcarrier spacing, one transmit and one receive antenna, an Open5GS core, and commercial handsets as UEs, three of them attached during the composition runs. Interface measurements come from the companion interface study [14], taken on the same host with the released mechanisms and raw percentiles under two conditions, a quiet host and the cell on the air (Figs. 12 and 13). The benchmark suite that produced them is published beside the platform. Tables VIII and IX separate the two kinds of evidence, mechanism costs and system-level runs, and say for each row what a reader needs in order to reproduce it. Limitations and status. The gaps are stated as plainly as the features, with what closes each.

about 100 µs of host jitter already means fronthaul lateness

rate floor: 1 CAS idle gate: 1 atomic load

Class B direct call: 0.29 µs P99.9

context cache read: 48 ns P50

1 slot (500 µs)

E3AP control RT: 11 µs P50

ring publish, 68 KB: 2.3 µs P99.9

1 ns 10 ns 100 ns 1 µs cost per event, log scale (measured on the GB10 host)

receiver kernels, live shape: 82 µs P50 (GPU, async)

10 µs

100 µs

1 ms

10 ms

UL-PHY worker, real-time priority, per cell: the only 4 operations it performs for dApps atomic gate 1 memcpy or 1 device-to-device enqueue try-lock + eventfd write CAS on the rate floor never on that thread allocation, locks that wait, any unbounded wait, ASN.1 or FlatBuffers encoding, log formatting, disk or network I/O; every residual wait is deadline-bounded and costs at most the slot management executors, normal priority: all the expensive work APER and FlatBuffers encoding, SCTP and socket I/O, lifecycle jobs, store leases, incidents, metrics, model staging Fig. 10. The latency architecture. Top: every steady-state mechanism on the real-time path on a log cost axis, three orders of magnitude inside the slot budget and far below the jitter that already causes fronthaul lateness. Bottom: work placement by thread. The real-time lane owns four cheap operations; everything expensive runs at normal priority.

author out-of-tree, SDK contracts, frozen C ABI

certify offline certify_class_a _cuda · L1 vectors · canary-checked outputs

package manifest-v1 · SPDX SBOM · detached sig · offline trust policy

admit hash/SBOM binding · isolated preflight · declared shape profiles reject

run guarded noexcept boundary · deadlines + breakers · incident records

fallback + breaker

isolate (Class C) SEQPACKET + creds · seccomp · rlimits · heartbeats · restarts kill + bounded restart

conventional PHY + scheduler, always armed: a fault, lateness, or unload returns stock behavior, never silence everything observable over E3: per-reason fallback counters, deadline misses, heartbeat latency, restart and incident records

Fig. 11. The safety ladder. Certification and signing happen offline; admission is fail-closed; runtime guards convert faults and lateness into fallback plus breaker; supervised isolation bounds even a misbehaving process. The conventional path remains in place at every rung, and every transition is observable over E3.

Class A tails. The reference receiver meets its budget at the live shape on an exclusive GPU, but the 273PRB kernels complete in about 270 µs against a 150 µs budget and are not yet latency-qualified. Until they are, admission profiles stop at the shapes the certifier has qualified, and a 100 MHz cell takes the conventional path for the wider grants. Tail qualification is the next admission gate. • Class B evidence was gathered on a host without approved CPU controls (no isolated cores or pinned realtime policy), so it is development evidence. The suite shows that on a quiet host every carrier meets the deadline and that only the direct call keeps its tail once the DU is on the air, which is the condition that matters. •

Steady-state overhead is guaranteed by construction and checked by fronthaul-lateness release gates, but the published comparison that the argument of Section XI calls for, the same traffic with the runtime disabled, enabled, and enabled with a full-rate recorder attached, reporting fronthaul lateness, GPU utilization, and per-core load, is still pending. It is the highest-priority measurement before this preview leaves the working group. • Over-the-air comparisons follow the discipline recorded in the operations cookbook, learned when a sequential comparison once ranked a no-op control above both real arms: interleave and reshuffle the arms, wait for link adaptation to settle after each switch, read CRC counters as deltas within a window from a single RNTI, and •

100 µs deadline

OCUDU direct ABI (validated call) SPSC queue, 2 threads

TABLE X P RIOR DA PP FRAMEWORKS ( BUILT TO PROVE THE TIER ) VERSUS THE OCUDU PLATFORM ( BUILT FOR PRODUCTION ADOPTION ). T HE DELTA IS SCOPE AND MATURITY, NOT INTENT.

SPSC shm, 2 processes SCTP loopback (E3-style RT) ZeroMQ REQ/REP ipc:// ZeroMQ REQ/REP tcp loopback 10−1

100

101

102

103

Prior (NEU / NVIDIA)

OCUDU dApp platform

Execution

External process over exported IQ/KPI

External and in-process; resident GPU L1 replacement; bounded scheduler hook

Control

Early primitives (e.g., PRB blacklist)

Typed, validated, allow-flag-bounded intents; granted quiet reservations; RAN control ops

Producer safety

Export coupled to L1

Never blocks; idle is one atomic; bounded waits; counted drops

Fallback

Left to the integration

Conventional path never displaced; breakers; quarantine

Packaging

Per-deployment integration

Signed manifest, SBOM, offline trust; preflight; ABI fingerprint

Lifecycle

Manual or scripted

E3 jobs, shadow state, generations, dry-run, rollback

Model ops

External to the framework

Stage, validate, warm, activate, rollback on a live cell

Multi-vendor

Vendor-neutral in intent; composition machinery still ahead

Runtime composition under operator trust and allow flags

Isolation

Container

Supervised process: authenticated IPC, seccomp, heartbeats, bounded restart

Wire

Pre-standard, still converging

Published ASN.1 APER and FlatBuffers, bounded codecs, versioned schemas; SCTP as first-class carrier

request → validated intents (µs); 20,000 requests per path P50 P99.9, quiet host

P99.9, quiet host, 50 µs antagonist on responder core P99.9, DU on the air on the same host

Fig. 12. The Class B contract under load. One 3.2 KB scheduler request over six boundaries, 20,000 requests each; bars are P99.9, ticks P50, the dashed line the admitted deadline. On a quiet host every carrier fits; with the DU on the air only the paths without a wake keep their tails. TABLE IX S YSTEM - LEVEL RUNS ON THE LIVE CELL . Run

Result

Conditions and gaps

Composition on one cell

four dApps, all three classes, more than 260,000 equalizer invocations, no fallbacks

out-of-tree neural equalizer (A), reference scheduler (B), spectrum and SRS-ISAC (C); three attached handsets. A control window with the same traffic and the dApps unloaded is not yet published two neural variants swapped by lifecycle alone, same UEs and traffic; CRC read per RNTI over activation windows. Transport-block counts, window lengths, and confidence intervals are not yet tabulated one E3 association; every state-changing tool dry-runs by default; reproducible against the loopback agent

Equalizer variants over the air

Operations surface

7 to 11 percentage points lower first-transmission BLER at MCS 11 to 15

eight telemetry streams; 28 MCP tools (twenty observe, eight act)

accumulate thousands of transport blocks per arm. The block counts, window lengths, and intervals behind the BLER row of Table IX and a control window for the composition run will be published with the qualification report. • Ecosystem. The multi-vendor machinery is implemented and exercised, yet every package that has run on the cell so far was produced by the authors’ organization. The out-of-tree neural equalizer demonstrates the seam, not yet an independent vendor. • Transport security. E3 carries no transport authentication in this release; the credentialed local socket, a loopback bind, and operator tunnelling are the boundary (Section VI). • Multi-cell. Telemetry attribution is single-cell today: every stream reports the publisher’s default cell id, and hostinline Class C is not a five-cell configuration. Multi-cell attribution and an asynchronous Class C prepared-output transaction are roadmap items. • MCP. The server is an example of the pattern, not a

hardened product. Its safety comes from the platform’s guardrails. XIII. W HAT THE E XPANDED C ONTRACT B UYS Table X summarizes the delta against the pioneering frameworks of Section II, and it should be read as a statement of scope and production maturity: those frameworks set out to prove the tier and succeeded, and much of what follows is machinery one needs only once dApps are trusted enough to matter. Three rows carry the argument. Authority: prior control returns were coarse primitives. Here every control path, Class A outputs, Class B intents, Class C maps and quiet requests, E3 RAN control, passes typed validation against operator-granted bounds, and the host remains the oracle. Producer safety: prior exports were coupled to L1; here the radio is arithmetically indifferent to its observers. Operations: prior dApps were hand-integrated deployments; these are governed artifacts with signatures, SBOMs, generations, dryruns, incidents, and rollback, which is what makes multivendor composition and agent-driven operation possible at all. XIV. C ONCLUSION AND AVAILABILITY Artifact availability. All software is open source under the BSD-3-Clause-Clear license in the OCUDU AI-RAN Working Group 2 group, https://gitlab.com/ocudu/work groups/wg2 ai ran/, as three repositories: ocudu-dapp-platform (the gNB with the dApp runtime, E3 agent, schemas,

103 102 10

what the dApp sees

P50 unsubscribed stream: 1 atomic load (P99.9) P99.9, prompt subscriber P99.9, subscriber sleeps 5 ms every 8th slot

1

100 10−1

OCUDU ring 8 slots

ZMQ PUB/SUB HWM 1000

age of data at the consumer (µs)

producer cost per slot (µs)

on the PHY thread

ZMQ PUB/SUB conflate

4457 dropped (counted) 0 dropped

2877 dropped (silent)

0 dropped

10588 dropped (silent)

13 dropped (silent)

105 104 ring depth × slot = 4 ms 103 102 101 100 OCUDU ring 8 slots

ZMQ PUB/SUB HWM 1000

ZMQ PUB/SUB conflate

Fig. 13. The producer rule, measured. Per-slot producer cost (left) and the age of the data the consumer sees (right) for the OCUDU shared ring and ZeroMQ publish/subscribe, with a prompt subscriber and one that sleeps 5 ms every eighth slot. The ring’s producer cost does not move, staleness is bounded by ring depth, and every drop is counted; ZeroMQ queues to its high-water mark or drops silently.

and Python client), ocudu-dapp-sdk (reference packages, certifier, scaffolds, examples, and the MCP server), and ocudu-dapp-quickstart (the reproducible container build, the loopback E3 agent, and the index of tutorials and documentation). The listings in Section VIII reproduce the E3 surface with no radio, GPU, or signing keys; the mechanism costs of Table VIII reproduce from the published benchmark suite on any host, with the CUDA host and the radio needed only where the table says so; the over-the-air results require a GB10-class host, a USRP or split-7.2 O-RU, and attached UEs, as the tutorials describe. This is a preview release: the working group invites feedback on the interfaces, new use cases that stress them, and independent vetting of the references, all of which feed the upstreaming of the runtime and E3 plane into the OCUDU mainline. Outlook. The ABI and schemas are frozen; the roadmap adds an asynchronous Class C prepared-output transaction, multi-cell telemetry attribution, and a broadening Class B feature block as evaluation-mode evidence accumulates. The platform is built to keep growing: new use-case interfaces, new streams, new references, and, by design, vendor dApps the authors will never see, composed by operators the vendors will never meet. Because the interfaces live in a vendor-neutral open-source project, the consensus they accumulate is portable into O-RAN and 3GPP when the time comes to formalize them. The dApp tier’s founders were right about where the opportunity lives, and the observer loop was the correct first step into it. This paper has presented the next step as a working system: an infrastructure on which the same governed artifact can observe the spectrum today, advise the scheduler tomorrow, and replace the receiver the day the evidence supports it, while feeding the datasets that train its own successor, answering to an operations agent, and sharing the cell with dApps from vendors who have never seen its code.

ACKNOWLEDGMENT This work was supported in part by the U.S. DoD OUSD(R&E) FutureG Office through the National Spectrum Consortium (NSC) Spectrum Forward OTA, including support for contributions to the Linux Foundation’s OCUDU opensource RAN project. R EFERENCES [1] M. Polese, L. Bonati, S. D’Oro, S. Basagni, and T. Melodia, “Understanding O-RAN: Architecture, interfaces, algorithms, security, and research challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 2, pp. 1376–1411, 2023. [2] O-RAN ALLIANCE, “WG3: Near-real-time RAN intelligent controller and E2 interface workgroup,” 2026, official workgroup charter. [Online]. Available: https://www.o-ran.org/technical-groups/wg3 [3] R. Wiesmayr, C. Dick, J. Hoydis, and S. Cammerer, “Design of a standard-compliant real-time neural receiver for 5G NR,” arXiv preprint arXiv:2409.02912, 2024. [4] S. D’Oro, M. Polese, L. Bonati, H. Cheng, and T. Melodia, “dApps: Distributed applications for real-time inference and control in O-RAN,” IEEE Communications Magazine, 2022, arXiv:2203.02370. [Online]. Available: https://arxiv.org/abs/2203.02370 [5] A. Lacava, L. Bonati, N. Mohamadi, R. Gangula, F. Kaltenberger, P. Johari, S. D’Oro, F. Cuomo, M. Polese, and T. Melodia, “dApps: Enabling real-time AI-based open RAN control,” Computer Networks, vol. 269, p. 111342, 2025, arXiv:2501.16502. [Online]. Available: https://arxiv.org/abs/2501.16502 [6] OpenRAN Gym and WiNES Lab, “dApp framework tutorial: Deploy and extend dApps,” 2026, openAirInterface spectrum-sharing tutorial. [Online]. Available: https://openrangym.com/tutorials/dapps-oai [7] WiNES Lab, “libe3: Vendor-neutral E3AP C++ library,” GitHub repository, 2026. [Online]. Available: https://github.com/wineslab/libe3 [8] D. Villa, M. Belgiovine, N. Hedberg, M. Polese, C. Dick, and T. Melodia, “Programmable and GPU-accelerated edge inference for real-time ISAC on NVIDIA Aerial testbed,” arXiv preprint arXiv:2512.06493, 2026. [Online]. Available: https://arxiv.org/abs/2512. 06493v2 [9] NVIDIA, “Aerial CUDA-accelerated RAN,” https://docs.nvidia.com/ aerial/, 2024. [10] K. Cohen-Arazi, M. Roe, Z. Hu, R. Chavan, A. Ptasznik, J. Lin, J. Morais, J. Boccuzzi, and T. Balercia, “NVIDIA AI Aerial: AI-native wireless communications,” arXiv preprint arXiv:2510.01533, 2025. [Online]. Available: https://arxiv.org/abs/2510.01533

[11] O-RAN Alliance nGRG, “dApps for real-time RAN control: Use cases and requirements,” O-RAN Alliance, Tech. Rep. RR-2024-10, Oct. 2024, contributed research report. [Online]. Available: https://www.o-ran.org/research-reports/ dapps-for-real-time-ran-control-use-cases-and-requirements [12] The Mobile Network, “When it comes to Nokia’s new AI-RAN, keep your eye on the dApps,” https://the-mobile-network.com/2026/ 07/when-it-comes-to-nokias-new-ai-ran-keep-your-eye-on-the-dapps/, Jul. 2026. [13] C. Boeira, E. Baena, A. Lacava, T. Melodia, D. Koutsonikolas, and I. Haque, “Performance characterization of dApps in open radio access networks,” arXiv preprint arXiv:2605.05426, 2026. [14] T. O’Shea, M. Pennybacker, and A. Kharchenko, “Enabling AI-RAN applications through dApp architecture decisions: A measured study,” 2026, companion paper, in preparation. [15] N. N. Santhi, D. Villa, M. Polese, and T. Melodia, “InterfO-RAN: Real-time in-band cellular uplink interference detection with GPUaccelerated dApps,” arXiv preprint arXiv:2507.23177, 2025. [Online]. Available: https://arxiv.org/abs/2507.23177 [16] M. Pennybacker, W. Liu, A. Kharchenko, and T. O’Shea, “GPUresident CUDA acceleration for OCUDU 5G PHY and O-RAN fronthaul: Architecture and preliminary performance,” arXiv preprint arXiv:2608.04338, 2026. [Online]. Available: https://arxiv.org/abs/2608. 04338 [17] OCUDU Ecosystem Foundation, “OCUDU ecosystem foundation FAQ,” https://ocudu.org/faq/, 2026, linux Foundation-hosted Open CU/DU ecosystem. [18] The Open Group, “POSIX dlopen().” [Online]. Available: https: //pubs.opengroup.org/onlinepubs/9699919799/functions/dlopen.html [19] NVIDIA, “CUDA runtime API: Stream management,” 2026. [Online]. Available: https://docs.nvidia.com/cuda/cuda-runtime-api/ group CUDART STREAM.html [20] ——, “CUDA programming guide: The CUDA driver API,” 2026. [Online]. Available: https://docs.nvidia.com/cuda/cuda-programming-guide/ 03-advanced/driver-api.html

Record · ID 667943 · SHA-256 6155efff4646a794
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.