1
Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
arXiv:2609.22753v1 [cs.DC] 19 Sep 2026
Delong Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu* Abstract—Natural-language service requests can require a language-model decision before execution starts, consuming part of the request’s latency budget. We integrate Jev’s decision-oriented application programming interface (API) into edge service orchestration to reduce this overhead while retaining service completion. The integration extracts four bounded intent fields and applies a shared validator, admission policy, and scheduler, accounting for decision waiting throughout the request timeline. We compare Jev with a short, structured-output DeepSeek deployment using live API measurements followed by modeled execution, and then a real two-node optical character recognition (OCR) service with self-hosted Qwen and rule-based references. Across three consecutive measurement blocks, Jev reduces median client decision latency by 15.9–26.5%. In eight paired OCR conditions, Jev matches DeepSeek’s correct, on-time completion count in seven and exceeds it in one. Without caching, median end-to-end latency on requests completed correctly by both systems is 11.1–25.3% lower; API fees per correct completion are 69.0–70.6% lower. Repeated-request caching largely removes the latency difference. The results demonstrate that decision-model substitution can lower both response latency and API fees in the tested service path, and identify fresh interpretation as the main opportunity for latency savings. Index Terms—Large language models, decision models, service orchestration, edge computing, service admission, quality of service.
✦
1
I NTRODUCTION
N
ATURAL - LANGUAGE interfaces let users describe an edge service and its execution requirements together: read the text in an image, keep the image at its originating site, and return the result before a deadline. Turning that request into an executable job requires both semantic interpretation and a placement decision. In edge computing, communication, resource availability, and placement jointly determine whether the service responds in time [1], [2], [3]. An interpreter on the admission path consumes part of that same response budget, delaying execution and leaving less time to complete the service. For a bounded service catalog, interpretation often requires only a small set of decisions, such as service type, placement permission, quality tier, and urgency. A large language model (LLM) can express these decisions through a short structured response, but the downstream scheduler needs the selected values rather than a generated explanation. A decision-oriented model could reduce this overhead even relative to a generative model configured for concise output. At the service level, the goal is to reduce this decision overhead while preserving correct, on-time completion under the request’s execution requirements. Intent-based networking already separates desired outcomes from their implementation [4]. Language-driven network systems and service prototypes connect interpretation to configuration and orchestration [5], [6], [7], [8], [9]. Small models and discriminative extraction offer further ways to specialize bounded language tasks [10], [11]. We
The authors are with the School of Electrical, Mechanical and Biomedical Engineering, University of Technology Sydney, Sydney, Australia. * Corresponding author: Guangsheng Yu (e-mail: [email protected]).
examine Jev, a decision-oriented service accessed through an application programming interface (API) [12], as the interpreter in an edge-service admission path. Queueing, communication, and execution all contribute to response time; their interaction with interpretation determines how much a faster decision benefits the service. We investigate three linked questions. How does Jev compare with a concise structured-output DeepSeek deployment in decision latency, semantic correctness, and billed API cost? Does the substitution lower full response time while retaining correct service completion? Under what workload conditions is faster interpretation most useful, and how does it interact with caching for repeated requests? We address these questions through live API measurements combined with modeled execution, followed by a real two-node optical character recognition (OCR) service. The real service transfers images, executes recognition, and checks the returned text against reference annotations. A fixed rule parser and self-hosted Qwen model extend the comparison beyond the two hosted APIs. The contributions are as follows. 1)
2)
We integrate Jev into an edge-service admission system through a four-field intent contract and a shared validator and scheduler. The system accounts for interpretation, queueing, communication, and execution on one timeline, connecting the choice of interpreter to service deadlines and completion. We measure a latency and API-cost benefit from this substitution. Jev reduces median decision latency by 15.9–26.5% across three blocks. In the real OCR service, it matches DeepSeek’s correct-completion count in seven paired conditions and exceeds it in
2
3)
one. Without caching, median full response time on requests completed correctly by both systems is 11.1–25.3% lower; API fees per correct completion are 69.0–70.6% lower across the eight conditions. We identify how request reuse changes the value of faster decisions. Applying the same cache policy to all backends shows that latency savings are concentrated in fresh interpretation, while repeatedrequest caching largely removes the difference. This comparison guides the use of fast decision models alongside caching in service orchestration.
2
R ELATED W ORK
2.1
Intent Interpretation and Service Orchestration
Intent-based networking separates desired outcomes from the mechanisms used to realize them [4]. Lumi translates natural-language network requirements through a learned interface and an intermediate representation [5]. NetConfEval evaluates language models on network-configuration tasks spanning policy translation, API calls, and configuration generation [6]; work on fifth-generation (5G) mobilenetwork intent extraction likewise studies the interpretation stage [13]. LLNet and fine-tuned small-model network configuration bring more specialized models into this setting [7], [14]. These studies establish natural-language interpretation as a systems component whose accuracy and execution overhead both matter. Several systems extend interpretation into service orchestration. Chat-Driven Optimal Management of Virtual Network Services combines an interpreter with optimization [8]. From Prompt to Service evaluates intent routing and real conversational service paths, including a finetuned small router [9]. Intent Engine uses intent grounding and validation within a network orchestration architecture [15]. Agentic infrastructure orchestration, zero-touch network pipelines, and DMO-GPT address broader automation workflows [16], [17], [18]. JAUNT studies intent-aware network tool routing [19], while grounded and role-based sixth-generation (6G) systems examine additional forms of orchestration and agent organization [20], [21]. Building on this work on service routing, we compare Jev with a generative interpreter while holding the intent contract and downstream policy fixed. We measure completion against the returned service output, accounting for decision waiting, case-sensitive OCR correctness, and caching under the same policy for each backend. The decision unit is a user-issued service job, whose admission path can include hundreds of milliseconds of interpretation. 2.2
Bounded Decisions and Small Models
Efficient prediction need not require free-form generation. Sentence-BERT and SetFit support lightweight semantic representations and classification [22], [23]. GLiNER and GLiNER2 provide discriminative extraction approaches, with GLiNER2 also addressing structured extraction and classification [11], [24]. The argument for small models in agentic systems further motivates assigning bounded tasks to specialized components [10]. For networking decisions, NetLLM adapts language-model representations to networking tasks through task-specific components [25].
TABLE 1 Shared intent contract. OCR denotes optical character recognition. Field
Allowed values
Service Locality Quality floor Urgency
Count, detection, OCR, unsupported Site only, remote allowed, unspecified Standard, high, unspecified Normal, urgent, unspecified
Alongside the hosted comparison, we include a fixed rule parser and an open generative model from the Qwen2.5 family [26]. Rules provide a low-overhead interpretation policy; Qwen provides a self-hosted generative alternative under the same service contract. 2.3
Serving, Structured Output, and Cost
Serving systems such as Clipper and InferLine address inference deployment under latency and resource constraints [27], [28]. PagedAttention and SGLang improve important aspects of language-model serving [29], [30]. Consequently, a measured self-hosted model latency is a property of its serving configuration as well as its weights. Structured-output validity also needs to be distinguished from task correctness. JSONSchemaBench evaluates efficiency, constraint coverage, and quality in structured generation [31]. A valid JavaScript Object Notation (JSON) object can still request the wrong service, just as a correctly interpreted request can yield incorrect OCR text. We retain these separate endpoints. Cost-aware prediction and routing systems, including FrugalML, FrugalGPT, and RouteLLM, consider different ways to allocate model calls under quality and cost objectives [32], [33], [34]. We compare fixed interpretation backends under a shared admission policy. Caching is another direct alternative to repeatedly paying inference latency. GPTCache studies reuse through a semantic cache [35]. Our cache reuses valid parsing results for identical normalized intent text. Applying the same policy to every backend tests how decision reuse changes the value of faster interpretation.
3
S ERVICE A DMISSION A RCHITECTURE
3.1
Request and Interpretation Contract
Each request i contains a natural-language description, an arrival time ai , an absolute deadline di , an origin, and payload metadata. Times share the originating controller’s clock. Only the description is sent to the interpreter; the scheduler receives the remaining metadata through the same interface for every backend. Reference intent labels and expected service outputs are available exclusively to the evaluator. Table 1 defines the shared interpretation contract. Its four fields describe the requested service and execution requirements, with 4 × 3 × 3 × 3 = 108 possible tuples. Missing requirements remain unspecified rather than being inferred from the evaluator’s labels. Jev selects among the allowed values, while the generative backends return the same fields in a concise structured response. The rule-based interpreter maps the same input text to this contract. The required meaning of each output is fixed across interfaces. A request to read an image locally
3
with high quality and urgent handling, for example, maps to OCR, site only, high, and urgent. Selecting the execution node remains the scheduler’s responsibility. 3.2
Validation, Caching, and Admission
The validator checks that all required fields are present and that each value belongs to its allowed set. Invalid interpretations do not proceed to scheduling. Requests also leave the admission path when their deadline or the available admission capacity is exhausted. This common policy controls which responses can be executed; semantic correctness must be evaluated separately against the request’s intended meaning. When caching is enabled, each backend reuses its own validated interpretation for identical normalized request text under the same extraction policy. A response becomes reusable only after it has arrived. The cache contains neither reference labels nor service outputs, and it can retain a wellformed but semantically incorrect interpretation. Repeated descriptions may accompany different images, so reusing a decision does not eliminate the corresponding service execution. All backends use the same cache policy. Admission and execution share one elapsed-time axis. Waiting for an interpretation slot and receiving a remote response both consume the budget between ai and di . A call that outlives its request’s deadline continues to occupy its slot until the call terminates, while new requests continue to arrive. Interpretation latency thus affects both service feasibility and queue occupancy. 3.3
Shared Scheduling and Execution
After interpretation, the scheduler reads the current worker state. Only an explicit remote permission allows a payload to leave its originating site; unspecified locality defaults to local execution. A high quality requirement selects the high tier, and other requests use the standard tier. Urgent requests take priority over ordinary requests without interrupting active work. These shared defaults can make distinct field values operationally equivalent, as with normal and unspecified urgency. For the real service, let ti denote the time at which request i is scheduled, after its interpretation has completed. Its selected tier is qi , and its priority is pi , with smaller values indicating higher priority. Let Ni contain the nodes that provide the interpreted service, satisfy its placement restriction, and have queue capacity. For a node n ∈ Ni , let bn be the estimated finish time of active work, or zero if idle, and let hn (q) be the calibrated duration of serving tier i q , including transfer and response overhead. Finally, Q≤p n is the set of pending jobs with equal or higher priority than request i, and qj is pending job j ’s tier. All node quantities refer to the state observed at ti . The predicted finish time fbin is X fbin = max(ti , bn ) + hn (qj ) + hn (qi ). (1) ≤p
j∈Qn i
The scheduler selects the node with the smallest predicted finish time among those satisfying fbin ≤ di , breaking ties in favor of the local node. It rejects the request if no such node
exists. Equation 1 is a shared admission estimate: executiontime variation and subsequent higher-priority arrivals can invalidate its prediction. Correct completion therefore depends on the observed result and finish time, not merely on passing the admission test. An accepted request transfers its payload to the selected worker and returns the service output to the originating controller. Payloads are transferred only after node selection. The evaluator checks the actual service, node, tier, priority, output, and deadline against the reference requirements. All elapsed times are measured at the originating controller. Locality constraints govern service payloads; the textual description still reaches the chosen interpretation API. The scheduler and these checks are common to all backends, allowing the comparison to focus on interpreter substitution.
4
E VALUATION M ETHODOLOGY
4.1
Comparison and Evidence Design
The principal comparison is between Jev and DeepSeek configured for concise structured output without a generated reasoning trace. Both use the same intent contract, admission limits, execution policy, and caching mechanism. The real-service comparison adds a self-hosted Qwen model and a fixed rule parser. Model identifiers, software, hardware, workload parameters, and execution procedures are provided in Appendix A. Study A measures live API decisions and combines their actual timing with modeled service execution. It evaluates the interpretation tradeoff and how decision waiting consumes execution slack. Study B executes a real two-node OCR service and checks the returned text. It tests whether lower decision latency remains useful after actual communication and processing. The measured hosted latency includes provider execution, routing, and network transport in both studies. 4.2
Workloads and Matched Comparisons
Study A evaluates 216 synthetic English requests covering the full intent contract. The same requests recur across three consecutive measurement blocks, with input order varied and backend call order alternated within pairs. Thus the blocks probe within-session variation on a shared corpus. An agent generates both the descriptions and their reference labels. Service conditions vary three factors: steady or bursty arrivals, changing or repeated descriptions, and caching disabled or enabled. Each of the eight conditions has a matched workload for the two hosted backends. In Study A, live admission events feed modeled execution on the same elapsed-time axis; no images are processed in that study. Study B holds the corresponding request descriptions, image assignments, arrival schedules, and service configuration fixed across all four backends. Complete backend runs are executed sequentially in randomized order. Consequently, matching controls the workload but does not eliminate variation in provider or network state. The real service uses Tesseract recognition [36] on images from the IIIT5K scene-text dataset [37]. Separate development images calibrate the scheduler’s service-duration
4
estimates. Evaluation images are selected before recognition results are observed and are not filtered by whether a tier recognizes them correctly. The tiers are execution configurations, rather than guarantees of recognition quality. Each condition includes 36 supported OCR requests and 12 unsupported requests. Images and intent descriptions are reused across conditions, while a new service run is still required after an interpretation-cache hit. Deadlines and offered loads are controlled study parameters. 4.3
Block
Backend
Correct
Median (ms)
p95 (ms)
1 1 2 2 3 3
Jev DeepSeek Jev DeepSeek Jev DeepSeek
214/216 216/216 213/216 215/216 212/216 216/216
319.5 434.7 314.7 425.9 320.7 381.4
529.3 835.9 557.8 833.2 624.2 686.9
Correctness and Completion
Exact semantic correctness requires all four interpreted fields to equal their reference values. Study A measures it on the complete paired request set. Operational completion instead requires a supported request to finish modeled execution by its deadline while satisfying the reference service, locality, minimum tier, and mapped priority. Its strict companion additionally requires exact field equality. Study B’s correct completion adds equality between the returned OCR text and the reference text after canonical Unicode normalization and removal of surrounding whitespace. Case and internal characters are retained. This distinguishes interpretation accuracy, execution compliance, and service-output correctness. For a given backend and condition, let K be the number of supported requests and let C be the number meeting the applicable completion criterion. For K > 0, the completion rate is C/K ; matched conditions share K , so their completion counts can also be compared directly. Unsupported requests and their correct rejections are reported separately. Outcome distributions over all arrivals retain both unsupported requests and failures. This prevents successful rejections from inflating service completion. 4.4
TABLE 2 Study A: exact four-field correctness and client decision latency. Each row contains 216 requests; the same texts recur across blocks.
Latency and API Cost
Decision latency is measured from client call initiation to receipt of the full interpretation response. Admission waiting is accounted for separately in the service timeline. We report medians and empirical 95th percentiles (p95). The figures also show empirical cumulative distribution functions (CDFs) and interquartile ranges (IQRs). These describe observed variation; they are not confidence intervals. For backend m and request i, let fim be the time when its service response reaches the originating controller. Its full request latency is Tim = fim − ai , including admission waiting, interpretation, service queueing, transfer, and execution. For a matched Jev–DeepSeek condition, let S contain the requests completed correctly by both backends, and let N = |S|. The median paired saving ∆paired is ∆paired = mediani∈S TiDeepSeek − TiJev . (2) Positive values favor Jev. The percentage reduction of median latency uses a different summary. Let Mm be the median of Tim over S ; the reduction RT , expressed as a percentage, is MJev RT = 100 1 − . (3) MDeepSeek
Equation 2 need not equal MDeepSeek − MJev . Both summaries are conditional on shared success, so we report N and each backend’s completion count alongside them. For a backend and condition, let A be the providerreported interpretation fees in United States dollars (USD) and let c be API cost per completion. Using the corresponding completion count C , we define
A , C > 0. (4) C The numerator includes all interpretation calls for that condition, including unsupported requests, unsuccessful jobs, and calls ending after their deadline. It excludes warmups and separate semantic measurements. Study A uses operational completion for C ; Study B uses correct OCR completion. Cost is undefined when C = 0. Relative cost reductions use the same ratio as Equation 3, with c replacing M . These fees exclude computation, communication infrastructure, energy, and maintenance. c=
4.5
Analysis Scope
The comparison units are matched conditions and consecutive measurement blocks. Repeated texts, images, and cached interpretations make arrivals dependent. We therefore report descriptive counts, quantiles, and differences. Observed retention means matching or exceeding the measured completion count under a matched condition. It does not establish statistical noninferiority: no applicationjustified noninferiority margin was specified.
5
R ESULTS
5.1
Decision Latency and Semantic Accuracy
Table 2 reports all three Study A blocks. Jev’s median full client decision time is 314.7–320.7 ms, compared with DeepSeek’s 381.4–434.7 ms. The corresponding reductions are 26.5%, 26.1%, and 15.9% in block order. Empirical p95 latency is also lower in each block, by 36.7%, 33.0%, and 9.1%, respectively. Jev’s advantage varies across the three blocks. Figure 1 shows the full distributions and their paired request-level differences. Both hosted APIs return schema-valid fields throughout Study A. DeepSeek is more accurate on the strict semantic endpoint: it gets 216, 215, and 216 requests correct, versus Jev’s 214, 213, and 212. Jev’s nine error occurrences span five distinct texts. Eight occurrences replace unspecified urgency with normal urgency; the ninth maps an unsupported recorded-speech translation request to OCR. This synthetic description combines an unsupported translation
5
Jev
(a) Decision-time distributions
DeepSeek
(b) Paired decision-time savings points: all pairs bar: IQR whiskers: 5–95%
1.0 Block 1
Empirical CDF
0.8 0.6
Block 2 0.4
216 responses / curve
Block 1
0.2
Block 2
Block 3
Block 3
0.0 200
500
1000
n = 216 matched texts per block
2000
4000
−1000
Client decision time (ms; log scale)
−100
0
100
1000 4000
DeepSeek − Jev (ms; symlog scale)
Fig. 1. Study A decision latency on the same 216 texts per model and consecutive block. (a) Full client response-time cumulative distribution functions (CDFs) on a logarithmic time axis. (b) All paired DeepSeek-minus-Jev differences: points, interquartile bars, nearest-rank 5th–95th percentile whiskers, and median diamonds. The symmetric-log axis is linear within ±50 ms. The ranges describe request variation, not confidence intervals; texts recur across blocks.
B1 · Jev
0
0
0
2
2
S/U/N
+6
+6
0
0
0
0
B1 · DS
0
0
0
0
0
S/U/C
0
0
0
0
0
-1
B2 · Jev
0
0
0
3
3
S/R/N
+1
+1
0
0
0
0
B2 · DS
0
1
0
0
1
S/R/C
0
0
0
0
0
0
B3 · Jev
1
0
0
3
4
B/U/N
0
0
0
0
0
0
B3 · DS
0
0
0
0
0
B/U/C
0
0
0
0
0
0
0
0
0
0
0
Locality
Quality
Urgency
Joint
B/R/N
0
Service
B/R/C
0
0
0
0
0
0
B1 Op
B1 Strict
B2 Op
B2 Strict
B3 Op
B3 Strict
Error occurrences / 216; joint = any field
Fig. 2. Study A semantic-error profile, with 216 responses per row. Joint counts indicate an error in any field; B1–B3 are consecutive blocks on reused texts, and DS denotes DeepSeek.
intent with an image-transfer permission clause. DeepSeek’s single error replaces remote permission with unspecified locality. Lower Jev decision latency therefore accompanies lower exact semantic accuracy. Figure 2 resolves the errors by field and block. Operational equivalence explains why this semantic gap need not reduce completion. Normal and unspecified urgency both map to ordinary priority in this scheduler, so that field mismatch can leave execution unchanged. In contrast, interpreting translation as OCR can dispatch the wrong service. The latter text was not selected into Study A’s execution traces, so its effect is absent from the modeled completion results. 5.2
Completion with Modeled Execution
Across Study A’s 24 paired conditions, Jev and DeepSeek have equal operational completion counts in 22, and Jev has higher counts in two (Fig. 3). Summing the fixed matrix yields 884/884 supported arrivals completed by Jev and
Jev − DeepSeek (completed requests / 48)
Fig. 3. Study A completion differences across all 24 paired conditions. Each cell is Jev minus DeepSeek, out of 48 arrivals. Op checks operational requirements; Strict also requires exact field equality. B1– B3 identify consecutive blocks. S/B denote steady/bursty arrivals, U/R changing/repeated text, and N/C cache off/on.
877/884 by DeepSeek. Each backend also receives 268 unsupported arrivals, which remain separate from execution successes. The strict companion endpoint gives 21 ties, two conditions higher for Jev, and one lower. In the latter condition, the loss is an unspecified-to-normal urgency mismatch with the same executed priority, in a condition with no cache hits. Both operational gains occur in the first block without caching. In its steady-changing condition, four DeepSeek calls take approximately 5.44–5.81 s, and four additional arrivals expire while waiting. Six of these eight failures are supported requests, yielding 36 Jev completions versus 30 DeepSeek completions. In steady-repeated traffic, one supported DeepSeek request lacks enough execution slack after waiting, yielding 41 versus 40 completions. These
6
(a) Correct OCR by condition
168
166
105
92
Correct rejection
92
80
26
96
20
20
Unsupported error
4
16
70
0
Qwen-7B
15
14
10
20
9
9
8
20
Deadline
0
4
99
0
Decision
0
0
0
96
Rules
10
10
13
13
10
10
13
13
Constraint
0
0
0
52
OCR error
120
118
84
48
Jev
DS
Qwen
Rules
Mutually exclusive counts / 384 per backend
B/
22
B/
22
B/
20
B/
20
S/
22
S/
20
S/
DeepSeek
S/
Correct OCR
20
R/ C
20
R/ N
22
U/ C
22
U/ N
20
R/ C
20
R/ N
22
U/ C
22
U/ N
Jev
(b) All-arrival outcomes
Correct / 48 arrivals (36 supported OCR)
Fig. 4. Study B real-service outcomes, one run per condition and backend. (a) Correct, on-time OCR completions in all 32 arms, each with 48 arrivals and 36 supported requests. (b) Mutually exclusive outcomes among all 384 arrivals per backend; DS denotes DeepSeek. S/B denote steady/bursty arrivals, U/R changing/repeated text, and N/C cache off/on.
TABLE 3 Study B: aggregate outcomes over eight conditions per backend. Completed OCR and correctly rejected requests have different denominators. False admissions count unsupported requests actually dispatched to OCR. Backend Jev DeepSeek Qwen2.5-7B Rules
Correct OCR
Correct rejection
False admit
168/288 166/288 105/288 92/288
92/96 80/96 26/96 96/96
4 13 4 0
trajectories show how delayed decisions can propagate into queueing and missed deadlines. 5.3
Correct Completion in the Real OCR Service
Table 3 aggregates the same eight conditions per backend. Jev completes 168/288 supported requests correctly and on time (58.3%), versus 166/288 for DeepSeek (57.6%). Qwen and rules complete 105/288 (36.5%) and 92/288 (31.9%), respectively. Figure 4 retains all 384 arrivals per backend, including the 96 unsupported requests. Jev matches DeepSeek’s correct-completion count in seven conditions and exceeds it by two in one (Table 4). Service-output errors remain substantial: the outcome breakdown attributes 120 Jev arrivals and 118 DeepSeek arrivals to OCR content or execution failure, after preceding failure categories have been excluded. Checking the returned text exposes these failures even when the decision and service calls themselves succeed. Among unsupported requests, Jev correctly rejects 92/96 and dispatches four unsupported requests to OCR. All four are repeated occurrences of one translation intent across conditions. DeepSeek correctly rejects 80/96, with 13 actual false admissions; the remaining three mishandled unsupported requests do not result in an OCR dispatch. Rules reject all 96 unsupported requests, but also lose supported jobs through interpretation and execution-constraint failures. No forbidden off-site image bytes are observed for any backend
in these runs, which include 22 remote executions for Jev and 30 for DeepSeek. 5.4
End-to-End Latency and the Cache Boundary
On the common-success subset, Jev reduces median full request latency by 11.1–25.3% in the four cache-disabled conditions. These comparisons contain 20 or 22 paired successful requests each. Median per-request savings range from 70.6 to 135.7 ms (Table 4). These durations include image transfer, worker queueing, OCR execution, and response delivery. Figure 5b retains every paired latency difference, including requests on which DeepSeek is faster. Figure 6 decomposes mean full time on the same successful subsets. Its service component includes image transfer and OCR; the remaining time includes controller handling and queueing. Caching changes the result when descriptions repeat. With caching enabled, the steady-repeated medians are 110.9 ms for Jev and 112.4 ms for DeepSeek; the burstyrepeated medians are 106.2 and 109.0 ms. Yet the median paired savings are −0.37 and −0.49 ms, respectively. The small opposite signs arise because a difference of marginal medians is not a median of paired differences. Neither comparison supports a material latency advantage after the reusable decisions are cached. Figure 5a also contrasts the uncached and repeated-text cached response-time distributions. Changing-text conditions have no exact cache hits; differences between their cache-on and cache-off arms reflect separate executions, not a benefit from cache lookup. Qwen has 105 correct completions in the full matrix and 99 arrivals classified as deadline or infeasibility failures. Its conditional median model-call time is 1.408 s, consuming a substantial part of the request budget before service execution. The called request subsets differ across backends because of cache hits and expiry, so this conditional median is diagnostic rather than a paired model-speed benchmark. With repeated-text caching, Qwen reaches 20 correct completions in both load conditions, matching the hosted backends there.
7
TABLE 4 Study B: all Jev–DeepSeek paired conditions. Each condition has 48 arrivals, of which 36 require OCR. Completion counts retain all failures; latencies use only the N requests completed correctly by both backends. Positive paired saving means faster completion with Jev. Correct completions Description
Cache
Jev
DeepSeek
N
Jev
DeepSeek
saving (ms)
Steady Steady Steady Steady Bursty Bursty Bursty Bursty
Changing Changing Repeated Repeated Changing Changing Repeated Repeated
Off On Off On Off On Off On
22 22 20 20 22 22 20 20
20 22 20 20 22 22 20 20
20 22 20 20 22 22 20 20
410.4 418.3 436.4 110.9 438.7 478.2 436.3 106.2
538.6 483.1 490.8 112.4 587.6 515.2 555.5 109.0
135.69 84.93 70.63 -0.37 101.14 74.43 125.86 -0.49
DeepSeek Mean full request time (ms)
1.0 0.8 0.6 0.4 Uncached (n=84) Cached repeat (n=40)
C R/ B/
C
N R/ B/
C
N
U/ B/
U/ B/
R/ S/
N R/ S/
Full request time (ms; log scale)
C
2000
N
1000
U/
500
250
S/
200
500
0
0.0 100
750
U/
0.2
Left: Jev Right, hatched: DeepSeek
Waiting / other Decision Service
1000
S/
Jev
60
Median paired
Load
(a) Full response-time distributions
Empirical CDF
Median full time (ms)
(b) Paired end-to-end savings points: all pairs bar: IQR whiskers: 5–95%
S/U/N · 20
Fig. 6. Study B mean time composition on each condition’s commonsuccess subset. The service component includes transfer and OCR; the residual includes queueing and controller handling. Left bars are Jev; right, hatched bars are DeepSeek. Components sum to the arithmetic mean, not the median. S/B denote steady/bursty arrivals, U/R changing/repeated text, and N/C cache off/on.
S/U/C · 22 S/R/N · 20 S/R/C · 20 B/U/N · 22 B/U/C · 22 B/R/N · 20 B/R/C · 20
−300 −100
0
100
300
1000
DeepSeek − Jev (ms; symlog scale)
Fig. 5. Study B latency on requests completed correctly by both Jev and DeepSeek. (a) Descriptive pooled CDFs; sample size n is 84 per model for the four uncached conditions and 40 for the two repeatedtext cached conditions. Changing-text cached arms are excluded from this panel. (b) Every paired difference in all eight conditions; row suffixes give n. Bars show the interquartile range, whiskers the nearest-rank 5th– 95th percentiles, and diamonds the median. The symmetric-log axis is linear within ±50 ms. Ranges are descriptive, not confidence intervals. S/B denote steady/bursty arrivals, U/R changing/repeated text, and N/C cache off/on.
5.5
Billed API Cost
Study A’s API fees per operational completion are 69.15– 72.57% lower with Jev across the 24 pairs (Fig. 7). These costs include unsuccessful requests and use observed modeled completion counts. The absolute fees and matched reduc-
tions retain all three blocks separately. Jev’s API fees per correct OCR completion are 68.97– 70.61% lower over all eight paired conditions; the cachedisabled range is 69.03–70.61%. Depending on the condition, the observed values scale to approximately $0.011–$0.067 per 1000 correct completions for Jev, versus $0.036–$0.215 for DeepSeek (Fig. 8). Caching reduces the absolute fee for both models and leaves a similar relative API-cost gap on the remaining calls.
6
D ISCUSSION
6.1
Implications for Service Orchestration
The experiments show a practical use for Jev in edgeservice admission: reducing the time and API fees spent interpreting requests while retaining the observed correctcompletion counts. The shared intent contract lets Jev replace the generative interpreter within the same validation and scheduling pipeline. The real OCR results show that the latency benefit remains visible after image transfer, worker queueing, and recognition. This integration gives service designers two ways to reduce interpretation overhead. A faster decision backend
8
Jev
effective the latter can be, including for self-hosted Qwen. Choosing between backends therefore depends on the share of requests that need a new decision and the time available for service execution. In the admission path studied here, shorter decision waiting leaves more deadline slack and releases interpretation capacity sooner.
DeepSeek
(a) API cost per completion S/U/N S/U/C S/R/N S/R/C
6.2
B/U/N B/U/C B/R/N B/R/C
3 offset pairs per row = 3 blocks
0.00
0.02
0.04
0.06
0.08
0.10
0.12
API USD / 1000 operational completions
(b) Matched cost reduction (%) S/U/N
72.05
69.33
69.36
S/U/C
69.25
69.33
69.36
S/R/N
69.95
69.19
69.42
S/R/C
69.20
69.15
69.26
B/U/N
69.25
69.33
69.36
B/U/C
69.25
69.33
69.36
B/R/N
69.20
69.19
69.42
B/R/C
72.57
69.15
69.19
Block 1
Block 2
Block 3
100 × (1 − Jev / DeepSeek); color range 68–74
Fig. 7. Study A billed API cost per modeled operational completion. (a) Absolute fees, scaled to 1000 completions; top-to-bottom offsets within each row distinguish blocks 1–3. (b) Matched reductions, with color range 68–74%. Fees include all trajectory calls and exclude warmups and separate quality evaluation. S/B denote steady/bursty arrivals, U/R changing/repeated text, and N/C cache off/on.
Jev
DeepSeek
Evaluating the Complete Service Path
The shared scheduler connects interpreter performance to delivered service. It reads current worker state after interpretation and checks placement, tier, priority, and predicted completion time before dispatch. Evaluating intent fidelity, execution compliance, and actual output separately makes it possible to trace how an interpretation affects the final result. For example, the normal-versus-unspecified urgency difference maps to the same execution priority, while the returned OCR text determines whether recognition succeeded. This evaluation supports a service-level choice of interpreter using completion, response time, and fees together. 6.3
Scope and Next Steps
The studies use synthetic English requests, repeated texts and images, and one measurement session; completion retention is descriptive as defined in Section 4. The real service has two workers and one service family. Hosted timings include provider and network paths, with sequential runs exposed to temporal variation. They compare deployed services rather than isolate internal model inference. API fees cover interpretation calls; local compute and communication costs are outside this measure. Appendix A records the configuration. The next step is to evaluate independently authored requests across days, services, and network conditions, using deadlines drawn from application requirements. Optimized local serving and a domain-trained classifier or bounded extractor, absent from the present comparison, would extend the available deployment choices.
S/U/N S/U/C
7
S/R/N
We integrated Jev into an edge-service admission pipeline and evaluated its ability to replace a concise generative interpreter. The shared contract and scheduler connect decision latency to actual service completion. Across three measurement blocks, Jev reduces median decision latency by 15.9–26.5%. In the real OCR service, it matches or exceeds DeepSeek’s correct-completion count in all eight paired conditions. Median response time among shared successes is 11.1–25.3% lower without caching. Across all eight conditions, API fees per correct completion are 69.0–70.6% lower. These results support Jev as a practical alternative for the tested admission workload. Faster decisions address requests that need fresh interpretation, while caching reuses decisions for repeated descriptions.
S/R/C B/U/N B/U/C B/R/N B/R/C 0.00
0.05
0.10
0.15
0.20
API USD / 1000 correct completions
Fig. 8. Study B API fees per correct OCR completion, scaled to 1000. Charges include all trajectory calls, including unsuccessful requests, but exclude warmups, computation, network, and other operating costs. S/B denote steady/bursty arrivals, U/R changing/repeated text, and N/C cache off/on.
shortens fresh calls, while caching reuses decisions for repeated descriptions. The repeated-text conditions show how
C ONCLUSION
R EFERENCES [1]
W. Shi, J. Cao, Q. Zhang, Y. Li, and L. Xu, “Edge computing: Vision and challenges,” IEEE Internet of Things Journal, vol. 3, no. 5, pp. 637–646, Oct. 2016.
9
[2]
Y. Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, “A survey on mobile edge computing: The communication perspective,” IEEE Communications Surveys & Tutorials, vol. 19, no. 4, pp. 2322–2358, 2017. [3] J. Santos, T. Wauters, B. Volckaert, and F. De Turck, “Towards lowlatency service delivery in a continuum of virtual resources: Stateof-the-art and research directions,” IEEE Communications Surveys & Tutorials, vol. 23, no. 4, pp. 2557–2589, 2021. [4] A. Clemm, L. Ciavaglia, L. Z. Granville, and J. Tantsura, “Intent-based networking - concepts and definitions,” RFC Editor, RFC 9315, Oct. 2022. [Online]. Available: https: //www.rfc-editor.org/rfc/rfc9315 [5] A. S. Jacobs, R. J. Pfitscher, R. H. Ribeiro, R. A. Ferreira, L. Z. Granville, W. Willinger, and S. G. Rao, “Hey, Lumi! using natural language for intent-based network management,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, Jul. 2021, pp. 625–639. [Online]. Available: https://www.usenix.org/conference/atc21/presentation/jacobs [6] C. Wang, M. Scazzariello, A. Farshin, S. Ferlin, D. Kostić, and M. Chiesa, “NetConfEval: Can LLMs facilitate network configuration?” Proceedings of the ACM on Networking, vol. 2, no. CoNEXT2, pp. 1–25, Jun. 2024. [7] A. Angi, A. Sacco, and G. Marchetto, “LLNet: An intent-driven approach to instructing softwarized network devices using a small language model,” IEEE Transactions on Network and Service Management, vol. 22, pp. 3403–3418, Aug. 2025. [8] Y. Miyaoka, M. Inoue, K. Urata, and S. Harada, “Chatdriven optimal management for virtual network services,” arXiv preprint arXiv:2512.24614, 2025. [Online]. Available: https://arxiv.org/abs/2512.24614 [9] L. Nisiotis and A. Hadjiliasi, “From prompt to service: An SLM-based agent orchestration gateway for AI-driven virtual worlds,” arXiv preprint arXiv:2606.03557, Sep. 2026. [Online]. Available: https://arxiv.org/abs/2606.03557 [10] P. Belcak, G. Heinrich, S. Diao, Y. Fu, X. Dong, S. Muralidharan, Y. C. Lin, and P. Molchanov, “Small language models are the future of agentic AI,” arXiv preprint arXiv:2506.02153, 2025. [Online]. Available: https://arxiv.org/abs/2506.02153 [11] U. Zaratiana, G. Pasternak, O. Boyd, G. Hurn-Maloney, and A. Lewis, “GLiNER2: Schema-driven multi-task learning for structured information extraction,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Suzhou, China: Association for Computational Linguistics, 2025, pp. 130–140. [12] D. A. Friedman, “Jev in practice: A composable Python toolkit for TypeSafe’s System One decision model,” Zenodo, Sep. 2026. [Online]. Available: https://zenodo.org/records/22817425 [13] D. M. Manias, A. Chouman, and A. Shami, “Towards intent-based network management: Large language models for intent extraction in 5G core networks,” in 2024 20th International Conference on the Design of Reliable Communication Networks (DRCN). Montreal, QC, Canada: IEEE, May 2024, pp. 1–6. [14] O. G. Lira, O. M. Caicedo, and N. L. S. Da Fonseca, “Network self-configuration based on fine-tuned small language models,” arXiv preprint arXiv:2512.02861, 2025. [Online]. Available: https://arxiv.org/abs/2512.02861 [15] K. Islam and R. N. Calheiros, “Intent Engine: Natural-language intent translation for intent-driven orchestration in the compute continuum,” Journal of Systems Architecture, vol. 179, p. 103938, Oct. 2026. [16] D. Brodimas, A. Birbas, D. Kapolos, and S. Denazis, “Intent-based infrastructure and service orchestration using agentic-AI,” IEEE Open Journal of the Communications Society, vol. 6, pp. 7150–7168, 2025. [17] D. Seo and K. Kim, “An integrated pipeline for intent-based zerotouch networks: From intent translation to minimal-modification reconfiguration,” Applied Sciences, vol. 16, p. 5811, Jun. 2026. [18] A. Mekrache, A. Ksentini, and C. Verikoukis, “DMO-GPT: An intent-driven framework for distributed 6G management and orchestration,” IEEE Communications Magazine, vol. 64, no. 1, pp. 48–54, Jan. 2026. [19] E. Li and H. Du, “JAUNT: Joint alignment of user intent and network state for QoE-centric LLM tool routing,” arXiv preprint arXiv:2510.18550, 2025. [Online]. Available: https: //arxiv.org/abs/2510.18550 [20] J. Martins, L. Mokrushin, M. Orlic, and A. K. A, “Intent-driven 6G service orchestration: Grounded translation, validation, and
decomposition,” arXiv preprint arXiv:2606.28348, Jun. 2026. [Online]. Available: https://arxiv.org/abs/2606.28348 [21] J. Parra-Ullauri, T. A. Khan, D. McHugh, S. Kapoor, A. Duke, A. Hey, and A. Corston-Petrie, “Role-based agentic AI for intent-driven network and service orchestration,” arXiv preprint arXiv:2606.20580, 2026. [Online]. Available: https: //arxiv.org/abs/2606.20580 [22] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China: Association for Computational Linguistics, 2019, pp. 3980–3990. [23] L. Tunstall, N. Reimers, U. E. S. Jo, L. Bates, D. Korat, M. Wasserblat, and O. Pereg, “Efficient few-shot learning without prompts,” arXiv preprint arXiv:2209.11055, 2022. [Online]. Available: https://arxiv.org/abs/2209.11055 [24] U. Zaratiana, N. Tomeh, P. Holat, and T. Charnois, “GLiNER: Generalist model for named entity recognition using bidirectional transformer,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Mexico City, Mexico: Association for Computational Linguistics, 2024, pp. 5364–5376. [25] D. Wu, X. Wang, Y. Qiao, Z. Wang, J. Jiang, S. Cui, and F. Wang, “NetLLM: Adapting large language models for networking,” in Proceedings of the ACM SIGCOMM 2024 Conference. Sydney NSW Australia: ACM, Aug. 2024, pp. 661–678. [26] Qwen, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu, “Qwen2.5 technical report,” arXiv preprint arXiv:2412.15115, 2024. [Online]. Available: https://arxiv.org/abs/2412.15115 [27] D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica, “Clipper: A low-latency online prediction serving system,” arXiv preprint arXiv:1612.03079, 2016. [Online]. Available: https://arxiv.org/abs/1612.03079 [28] D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. Gonzalez, and A. Tumanov, “InferLine: Latency-aware provisioning and scaling for prediction serving pipelines,” in Proceedings of the 11th ACM Symposium on Cloud Computing. Virtual Event USA: ACM, Oct. 2020, pp. 477–491. [29] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in Proceedings of the 29th Symposium on Operating Systems Principles. Koblenz Germany: ACM, Oct. 2023, pp. 611–626. [30] L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng, “SGLang: Efficient execution of structured language model programs,” arXiv preprint arXiv:2312.07104, 2023. [Online]. Available: https://arxiv.org/abs/2312.07104 [31] S. Geng, H. Cooper, M. Moskal, S. Jenkins, J. Berman, N. Ranchin, R. West, E. Horvitz, and H. Nori, “JSONSchemaBench: A rigorous benchmark of structured outputs for language models,” arXiv preprint arXiv:2501.10868, 2025. [Online]. Available: https://arxiv.org/abs/2501.10868 [32] L. Chen, M. Zaharia, and J. Zou, “FrugalML: How to use ML prediction APIs more accurately and cheaply,” arXiv preprint arXiv:2006.07512, 2020. [Online]. Available: https://arxiv.org/abs/2006.07512 [33] ——, “FrugalGPT: How to use large language models while reducing cost and improving performance,” arXiv preprint arXiv:2305.05176, 2023. [Online]. Available: https: //arxiv.org/abs/2305.05176 [34] I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs with preference data,” arXiv preprint arXiv:2406.18665, 2024. [Online]. Available: https://arxiv.org/abs/2406.18665 [35] F. Bang, “GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings,” in Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023). Singapore, Singapore: Association for Computational Linguistics, 2023, pp. 212–218.
10
[36] R. Smith, “An overview of the Tesseract OCR engine,” in Ninth International Conference on Document Analysis and Recognition (ICDAR 2007) Vol 2. Curitiba, Parana, Brazil: IEEE, Sep. 2007, pp. 629–633. [37] A. Mishra, K. Alahari, and C. Jawahar, “Scene text recognition using higher order language priors,” in Proceedings of the British Machine Vision Conference. Surrey: British Machine Vision Association, 2012, pp. 127.1–127.11.
A PPENDIX A I MPLEMENTATION AND E VALUATION S ETTINGS A.1
Interpretation Backends
The studies were executed on September 18, 2026. Jev uses OpenRouter’s Decisions endpoint with identifier typesafe/jev-1.13; responses identify jev-1.13-20260917 and provider TypeSafe. Four native Choice questions return the shared intent fields in one request. DeepSeek uses deepseekv4.1-flash through OpenRouter, pinned to Together with provider fallback disabled. It returns a strict four-field JSON object with temperature zero, a 128-token limit, and reasoning disabled; recorded reasoning-token counts are zero. Qwen2.5-7B-Instruct, revision a09a35458c70, runs in brain floating point 16 (BF16) on one NVIDIA L40, using Transformers with scaled dot-product attention (SDPA) and PyTorch 2.2.0+cu121. It uses the same generative prompt and field contract as DeepSeek, greedy decoding, and a 128token limit. The server batches up to four already-pending requests without an intentional collection delay. The fixed rule parser uses keywords and regular expressions without tuning on evaluation wording. A.2
Admission and Execution Controls
Both studies use a 2 s request deadline, four concurrent interpretation slots, and an admission queue of capacity 32. API calls have a 15 s timeout, with no automatic retries or repair calls. Connections are warmed before measurement. Each condition starts with an empty cache keyed by normalized full request text and extraction policy. The controller checks it before queueing and after obtaining a slot; concurrent misses are not coalesced. The real service has one local macOS worker and one remote Linux worker, connected through a Secure Shell (SSH) tunnel. The remote allocation provides six processor cores and the accelerator used by Qwen; OCR uses processors only. Each node has one nonpreemptive worker and a queue of capacity 16. Urgent and ordinary jobs have priority values zero and one, respectively; equal-priority jobs follow enqueue order. The controller verifies the returned node, tier, priority, and image identity. A timed-out OCR process is terminated before capacity is released. Transport or protocol errors stop further calls. The recorded environment lacks a complete processor inventory for exact cross-site reproduction. A.3
Study A Parameters
The 216 requests use six wording families and two instances of each of the 108 field combinations, with extraction instructions held fixed. Seeds 42, 43, and 44 vary request order and arrival traces across three consecutive blocks. Each model has 648 semantic measurements on the same
216 texts. There is one semantic-stage warmup per model per block and no fixed interval between blocks. Each service condition has 48 arrivals. Steady traffic has rate 2 requests/s; bursty traffic alternates 0.5 and 8 requests/s in 20 s segments. The execution model has four edge nodes and one cloud node, with payloads of 0.25– 2 MB. Local, other-edge, and cloud links have propagation delays of 2, 20, and 60 ms and bandwidths of 1000, 100, and 50 Mbit/s, respectively; transfer also includes serialization. Base durations are 40 ms for counting, 80 ms for detection, and 60 ms for OCR. The high tier multiplies these by 1.8; node factors are 1.0, 1.1, 1.2, and 1.3 for the edges and 0.65 for the cloud. Each node has one priority worker. These are exploratory simulation inputs. A.4
Study B Parameters and Scoring
Model conditions run sequentially in randomized order with seed 42 and no fixed cooldown. Both workers use Tesseract 5.5.3, English recognition, one OpenMP thread, and single-word mode (engine mode 1, segmentation mode 8). Standard and high tiers use tessdata fast and tessdata best, respectively, with identical weights across nodes. Seed 42 selects 20 training and 80 test images from sorted IIIT5K identifiers before recognition outcomes are known. Calibration runs each training image once per node and tier. Median request–response durations, including transfer, are 80.85/96.26 ms for local standard/high and 95.30/139.35 ms for remote standard/high. The workload has 48 distinct intents. Repeated-text conditions use eight descriptions, six supported and two unsupported, each repeated six times. The test images recur across conditions. Arrival rates match Study A; eight conditions for four backends produce 32 runs and 1536 arrivals. OCR comparison uses case-preserving character annotations without a recognition lexicon, normalizes Unicode to Normalization Form C (NFC), and strips surrounding whitespace. Quantile summaries use nearest-rank 5th and 95th percentiles and linearly interpolated interquartile ranges. Each request is evaluated on its recorded timing and output; cached decisions do not replace service execution.