Conceptio › Archive › arXiv CS
arXiv CSopen access

Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache Frank Li

arXiv:2609.15030v1 [cs.DC] 14 Sep 2026

UNSW Sydney ・[email protected] Abstract. External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache under four-way tensor parallelism. A complete-hit recovery mismatch restored state for the full prompt while the scheduler credited one fewer token. We aligned recovery through strict-prefix lookup and established a numerical comparison using shared computation corrections, matched checkpoint scheduling, and fixed per-rank kernel configurations. In a nine-length serial workload, agreement with the modified recomputation control improved from 34/36 to 36/36 generations, each containing 64 token IDs. A separate instrumented run passed recorded transfer-page, effective-tail, and delayed-save checks. Three additional synthetic templates passed 72 paired 256-token continuations across two fresh-container runs. A subsequent serial performance study preserved output equality across 120 requests; among the measured trials, CPU reload reduced time to first token by 46–64% and total request time by 1.9–7.0% relative to modified cold recomputation. The contribution is an experimentally validated integration repair applying an existing checkpoint-alignment principle. The evidence is confined to one model revision and controlled configuration; it does not establish general determinism, task-quality equivalence, concurrent-serving gains, or capacity beyond GPU memory.

1 Introduction

treatment. The main evaluation uses the final, explicitly identified comparisons.

At a prompt length of 3,584 tokens, four transfer workers reported restoring a cached checkpoint for all 3,584 tokens. The scheduler, however, credited only 3,583 tokens and scheduled one additional input token. The subsequent continuation differed from the no-connector control at generated-token index 11. Transfer had occurred on every tensor-parallel rank; a positive cache-hit counter did not explain whether computation resumed from the correct logical position. This distinction matters when a cache contains state updated in place as well as attention history. Reducing a token count cannot reconstruct an earlier recurrent state. Existing work already describes the need to align reusable attention prefixes and recurrent checkpoints; we apply that principle to a concrete integration failure rather than propose a new checkpointing algorithm [1]. The failure was difficult to isolate because the output reference also required validation. Earlier paired runs could diverge despite identical recorded weights, and apparently successful byte checks had initially inspected ineffective all-zero tail locations. Neither source inspection nor one successful transfer was sufficient evidence. The eventual comparison therefore used a modified recomputation control with shared computation changes, a matched checkpoint schedule, and a fixed per-rank Flash Linear Attention (FLA) kernel profile. It did not compare a connector-only patch against an untouched production executable. This case study addresses three questions: which recovery condition failed in the integration; how a useful numerical control was established; and what the repaired path demonstrably passed. Its contributions are an observed complete-hit state/position mismatch and its repair, a documented construction of the numerical comparison, and a bounded regression combining generatedtoken agreement with transfer and completion checks. The historical investigation contains 53 changing experimental rounds, not 53 independent repetitions of one

2

System and comparison design

2.1

Execution and cached state

The experiment runs one full GLM-5.3-Flash target model across four GPUs with tensor parallelism. GLM5.3-Flash combines sparse and linear attention; the evaluated NVFP4 checkpoint is Red Hat’s quantized derivative of the Z.ai model. Our workloads use text inputs only [8], [9]. Throughout this paper, “GLM” abbreviates this evaluated model and checkpoint, not the GLM model family as a whole. vLLM performs model execution and scheduling; LMCache provides the external cache path. The four ranks are parts of one inference engine, not four independent model replicas. The experiment does not implement prefill/decode disaggregation. Figure 1 summarizes the execution and recovery responsibilities. The integration handles attention-related cache entries, checkpointed state, and model-specific tail entries. Logical engine groups and physical transfer groups are distinct: the audited registration contains 70 entries per rank across six transfer groups. Some transfer representations are byte views, while the observed tail entries are BF16. Thus neither the model checkpoint’s NVFP4 name nor the FP8 KV configuration describes every state tensor’s precision. MTP configuration does not imply that every step drafts a token: checkpoint-protection steps can disable drafting. The later content extension records four NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs and their UUIDs. This current inventory does not retroactively establish the exact devices used in earlier rounds. Performance-specific runtime settings are reported in §5.5; the shared computation controls are described below.

1

Shared controls: computation corrections · matched checkpoint schedule Per-rank FLA profile Save / restore

One GLM-5.3-Flash engine TP = 4

vLLM scheduler Computed prefix p

R0

R1

R2

LMCache connector

R3

External CPU cache

Attention cache + checkpoint state + tails

External path enabled only in candidate Logical components: four ranks belong to one engine; hardware interconnects are not depicted.

Figure 1: Evaluated logical components. The four tensor-parallel ranks form one GLM-5.3-Flash engine. Both arms share modified computation controls; only the candidate enables the external CPU-cache path. Arrows describe logical interactions, not hardware topology. Table 1: Evaluated configuration Setting

Evaluated configuration

Target Model revision Runtime Parallelism Precision Execution Limits

Full 45-layer RedHatAI/GLM-5.3-Flash-NVFP4

Checkpoint interval Numerical control

36c184c6cda000a481711306df5adde42f63321a

vLLM 0.1.dev20051+g487ecf187; LMCache 0.5.4; PyTorch 2.13.0+cu130 TP=4, PP=1, DP=1; four GPUs on one host Configured model dtype: BF16; compressed-tensors quantization; FP8 KV setting Eager, synchronous scheduling; MTP configured with one speculative token 4 GiB KV budget per rank; maximum model length 32,768; maximum sequences 4; token batch limit 8,192 C=1,792 tokens Seven recorded FLA autotuner configurations per rank

2.2 State position and scheduler position

their stated combination; it does not establish that the cumulative patch alone preserves production outputs. The final strict-prefix intervention changes two GLM lookup entry points within this existing configuration.

Let N denote prompt length, C the checkpoint interval, B the logical prefix represented by the restored checkpoint, and p the scheduler’s computed-token count. Prefix lengths are counts; token indices are zero based. A state representing prefix [0,B) supports continuation from B. Crediting p=B therefore aligns the restored state with the start of subsequent computation. For the fully cached prompts in the final workload, restricting lookup to the prompt without its last token selects B=C floor((N−1)/C). The scheduler then computes the nonempty suffix [B,N). This formula assumes the aligned checkpoint is available, as it was in these reload tests. In a general cache, lookup must select an available checkpoint within the query boundary; the formula is not a guarantee of a hit.

3

Integration failures and repair

3.1

Observing effective state and waiting for saves

Early tail checks reported equal all-zero pages. Later examination of the effective locations and strides invalidated that coverage claim: equality at an irrelevant address was not a test of the state consumed by subsequent computation. Correcting the addressing exposed nonzero tail data and transfer discrepancies that the earlier observation had missed. The earlier success interpretation was retained as a corrected failure of observation, not counted toward the final validation. Save submission also had to be distinguished from completion. The integration was changed to wait on the actual DeviceMessagingFuture before allowing dependent state reuse. The final regression inserted a 500 ms delay into saves and paired the actual waits with the delayed operations. This tests whether the recorded dependency is respected under the injected condition; it does not estimate a production failure probability.

2.3 Treatments and observations The cumulative functional patch changes eleven files across LMCache and vLLM. It includes layout and snapshot handling, save completion, checkpoint behavior, and stable selection/packing changes. The FLA installer and the no-connector scheduling fallback are separate controls. Consequently, successful evaluation is a result for

2

Prompt length N = 3,584; checkpoint interval C = 1,792

Restore state B = 3,584

Before

Scheduler credit p = 3,583

Compute 1 input token First difference: output index 11

B ≠ p: changing a count does not roll back loaded state Restore state B = 1,792

After

Scheduler credit p = 1,792

Compute 1,792 input tokens All 64 output tokens agree

B = p: look up a strict prefix, then resume at the matching position

Figure 2: Checkpoint alignment at N=3,584 and C=1,792. Before repair, the loaded state represents B=3,584 while the scheduler credits p=3,583. Strict-prefix lookup instead restores B=p=1,792 and recomputes the remaining suffix. This is a logical boundary example, not a timing trace. Table 2: Treatments and observation conditions Configuration

Computation and scheduling

Connector

Observation conditions

Original production configuration Modified recomputation control

Original production settings

Not the evaluated treatment Disabled

Not the output control in the final matrix CPU batch/configuration logging; no GPU tensor probes

Modified candidate

Instrumented candidate

Shared computation corrections, fixed perrank FLA profile, matched checkpoint schedule Corresponding common corrections and Enabled; lookup FLA profile; connector checkpoint path boundary changed between the two matrices Repaired candidate configuration Enabled

These checks answer separate questions. Comparing source and destination pages checks the audited bytes. Observing nonzero effective tails checks that selected useful state entries are exercised. Waiting for the actual save checks a completion dependency. None alone establishes that the bytes describe the prefix from which the scheduler resumes.

CPU batch/configuration and transfer logging; no GPU tensor probes GPU byte checks, destination poisoning, tail checks, injected save delay

an earlier complete checkpoint for an exact-boundary prompt. Restricting the query before checkpoint selection also keeps selection and its associated lookup bookkeeping at the same boundary; subtracting a count after retrieving full state would leave the original problem intact. At N=3,584, the repaired path loads B=1,792 and computes a 1,792-token suffix. At N=5,376, it loads B=3,584 and computes the same suffix length. Both previously failing continuations then match their controls. This repair increases recomputation relative to the faulty onetoken schedule at exact boundaries. The incremental latency cost relative to that incorrect path has not been measured; §5.5 instead compares repaired reload with modified cold recomputation.

3.2 Complete-hit boundary mismatch The fixed-profile pre-fix matrix exposed failures at N=3,584 and N=5,376. For both lengths, raw transfer events from every rank reported a checkpoint covering the entire prompt, B=N. The scheduler’s computed count and external-hit metric instead reported p=N−1, and the first reload batch scheduled one input token. The archived connector explains the discrepancy. It retained the aligned lookup result in the request tracker used for snapshot retrieval. A later complete-prompt branch reduced the amount credited to the scheduler by one. Snapshot selection still used the retained full boundary. The observed condition was therefore B=N and p=N−1, rather than a snapshot of the shorter prefix. This mechanism is supported by the actual transfer and batch records, not inferred from the metric alone. The repair excludes the final known token before lookup at both GLM entry points. Alignment then selects

4

Establishing a numerical control

4.1

What the diagnostic comparisons established

Temperature zero and a fixed seed did not make the original cross-process comparison a reliable attribution test. Recorded batches differed: an earlier no-connector path processed a 3,583-token prefill, while the connector path split it into 1,792 and 1,791 tokens; decode scheduling differed as well. A matched checkpoint schedule removed this particular discrepancy, but output differences remained.

3

A subsequent audit found matching named parameter and buffer storages across arms. That narrowed the investigation without proving that all scratch allocations or execution choices were identical. An actual mHC invocation then showed that the attention-output input already differed between arms, while the other seven captured tensor inputs and scalar arguments agreed. This observation does not demonstrate a same-input mHC defect. Probes taken in separate runs also cannot be joined into one observed layer-by-layer causal trace. The relevant prior distinction is between repeated execution at one shape and invariance across batching or prefix splits. Fixed seeds do not resolve changes in floatingpoint reduction geometry. He discusses this distinction and concrete kernel controls in an author technical report; those results motivate the local control but do not validate this GLM implementation [2].

request, with temperature zero, seed 42, and EOS ignored. No recorded output contains the known EOS token IDs. Although the prompt requests a longer explanation, the 64-token cap is a finite continuation test, not evidence of satisfying that instruction or preserving task quality. We distinguish paired output equality, within-engine cold/reload equality, actual four-rank transfers, and internal transfer observations. There are 36 paired comparisons in each two-arm matrix, not 36 distinct tasks. Preemption counters remain zero. The serial reset/interposer procedure exercises the cache path; it does not demonstrate behavior under sustained memory pressure or concurrent preemption. Evidence checks separately recompute results from archived records and cross-check recorded observations; they do not provide an independently implemented reference for every low-level checker. Hash agreement establishes byte identity, not correct state addressing. The retrospective page audit verifies recorded comparisons rather than re-reading historical device buffers.

4.2 Recorded FLA profile and its scope The experiment recorded actual FLA autotuner choices in a baseline run and pinned the seven observed configurations separately for each rank. The installer constrained each observed tuner to its recorded configuration and cleared its selection cache. Invocation of an unrecorded tuner was rejected. Later runs additionally checked the actual selected configuration after calls. The profile was a joint intervention on seven configurations per rank. No single-kernel ablation establishes that one tile size, warp count, or kernel alone caused the earlier divergence. The profile is neither a general determinism algorithm nor an established performance optimum. An older reference from a different diagnostic round remained unequal under the fixed profile. Its failed assertion and nonzero process exit were preserved. The selected profile reference was identified before the intervention; agreement with it does not turn the older comparison into a pass. This distinction prevents a changing numerical reference from silently changing the reported acceptance criterion.

5.2

Boundary results without GPU tensor probes

Table 4 retains all nine cases. “First difference” is the zero-based generated-token index in the pre-fix reload comparison; a dash means the complete 64-token continuation agrees. B comes from transferred checkpoint metadata and p from the first reload batch. Each repaired row contains four paired generations: target cold, two interposers, and target reload. The pre-fix matrix has 34/36 equal generation pairs, with the two differences confined to the exact-boundary reloads. The repaired matrix has 36/36 equal pairs, and all nine target cold/reload comparisons agree. All 36 control generations also agree between the pre-fix and repaired runs. Across the repaired matrix’s nine reloads and four ranks, the first batch has computed=B and scheduled=N−B. Every reload records positive H2D transfer on all four ranks. At N=3,584, transferred bytes per rank change from 69,625,856 to 54,809,600; at N=5,376, they change from 84,442,112 to 69,625,856. These metadata corroborate the change in selected checkpoints. These before/after byte counts alone do not establish a latency benefit. Section 5.5 separately measures the repaired path against modified cold recomputation; it is not a performance comparison against the incorrect pre-fix path. An earlier attempt to execute the matrix inadvertently ran only the five-operation first case because of a harness import entry point. The independent coverage verifier rejected it. It is excluded from the nine-case results. Likewise, the changing diagnostic rounds are not pooled into a statistical reliability estimate.

5 Evaluation 5.1 Workload and metrics The same archived plan is used for the pre-fix, repaired, and instrumented matrices. It contains nine prompt lengths around selected boundaries: 3,583, 3,584, 3,585, 3,587, 3,588, 3,589, 5,375, 5,376, and 5,377. Each case performs five serial operations: generate the target from a cold state, generate two interposer requests, reset the GPU cache while retaining the external cache, and regenerate the target. Each paired arm therefore executes 45 operations, including 36 generations and nine resets. The inputs comprise 27 distinct token sequences generated from one Chinese password-retrieval and explanation template. Target and interposer variants differ in a document tag near the start; they are not independent application domains. There is no explicit request cache salt, and recorded transfer events have an empty salt. The output criterion compares all 64 generated token IDs per

5.3

Instrumented transfer and save regression

A separate repaired candidate executes the complete plan with byte observers, destination poisoning before H2D, and injected save delay. Its 36 generations match the archived repaired-run control, and all nine cold/reload pairs agree. This is an instrumented candidate compared

4

Table 3: Numerical-control experiments Experiment

Observation

Interpretation

Profile capture

Recorded per-rank configurations and actual call inputs

A concrete reference for the next intervention

Fixed-profile first case

All four 64-token generations agree between arms and with the selected reference

Supports the controlled paired comparison

Fresh-container repeat

New engines reproduce those first-case comparisons

Independent repetition of that bounded test

Full boundary matrices

Actual configurations verified in both arms on four ranks

Numerical control retained during the boundary intervention

Table 4: Complete boundary matrix before and after repair N

B before

p before

First difference before B = p after

Suffix after

Equal generations after

3583 3584 3585 3587 3588 3589 5375 5376 5377

1792 3584 3584 3584 3584 3584 3584 5376 5376

1792 3583 3584 3584 3584 3584 3584 5375 5376

— 11 — — — — — 0 —

1791 1792 1 3 4 5 1791 1792 1

4/4 4/4 4/4 4/4 4/4 4/4 4/4 4/4 4/4

1792 1792 3584 3584 3584 3584 3584 3584 5376

with a previously recorded control, not a new two-arm run. Page rows are repeated internal checks, not 20,928 independent round trips or randomized trials. Nonzero tails demonstrate coverage of the selected entries, not equality of every internal model state. The artificial delay and heavy observation deliberately change execution conditions, so this run is not a serving-performance measurement. Its agreement complements the output matrix without GPU tensor probes; it does not replace that matrix.

zero draft tokens in the protected first batch. Actual FLA profiles match the recorded policy; no known EOS token or request preemption is observed. Each run exports 49 files whose SHA hashes are independently verified. This adds 72 successful 256-token generation pairs and eight repeated bridge pairs across two container-level repetitions. These are repeated synthetic inputs, not 80 independent tasks. A setup attempt was aborted before workload execution after detecting an incorrect served-model identifier in the new harness; it has no model-correctness verdict and is excluded. The prompt contents, continuation lengths, output reference, and acceptance criteria were not changed in response to model results. This extension does not repeat the poisoned byte audit, establish task-answer quality, or measure performance. The complete plan, failed setup record, raw exports, and independent checks are retained in the internal evidence archive.

5.4 Additional prompt templates and longer continuations A subsequent preregistered extension tests three additional synthetic templates: English ledger arithmetic, Chinese ordered-rule application, and English Python-code reasoning. Each template uses prompt lengths 3,583, 3,584, and 5,376, with the same cold/interposer/reset/reload procedure and 256 generated tokens per new request. The historical first case remains as a 64-token bridge to the previously selected reference. The combined plan contains 30 distinct input sequences and 50 operations per arm: four bridge generations and 36 extended generations, totaling 9,472 output tokens per arm. Two separate disposable containers each run a fresh no-connector engine and a fresh candidate engine under the existing fixed configuration. The exported token plans are identical. Both pairs pass all 40 continuation comparisons, and the two runs also agree on every baseline and candidate continuation. For all ten reloads per run, fourrank transfer metadata and first-batch records agree on B=C floor((N−1)/C), computed=B, scheduled=N−B, and

5.5

Paired serial performance

A separately preregistered study measures serial streaming completions on the disposable container’s loopback interface. Two fresh containers run the same modified controls in opposite orders: OFF→ON, then ON→OFF. Each engine runs one excluded warmup and three measured trials per length, N∈{3583,3584,5376}, with distinct early document identities and 128-token continuations. OFF attempts cold and immediate-repeat requests; ON additionally attempts CPU reload after a GPU-only reset. All 120 requests (90 measured) pass complete output-token checks: each input has the same output across its five request conditions and both containers. The four engines are two paired repetitions, not 120 independent experiments. The workload uses one English 5

Table 5: Instrumented transfer and save regression Check

Observed coverage

Result

Candidate/control generations

36 continuations × 64 tokens

All agree

D2H page-comparison rows

17,640

Zero reported byte mismatches

H2D page-comparison rows

3,288

Zero reported byte mismatches; destinations poisoned

Combined page-comparison rows

20,928

All 70 registered entries covered per rank and direction

Nonzero tail H2D observations

4 ranks × 9 reloads × 12 entries = 432

All observed entries nonzero

Delayed-save wait pairs

67 per rank; 268 total

Actual waits cover the injected 500 ms delay

H2D volume

2,447,265,792 bytes across four ranks

Positive transfers for every reload

Table 6: Content and continuation extension across fresh engines Measure

First pair

Fresh repeat

Operations per arm Historical 64-token bridge comparisons New 256-token comparisons Output tokens per arm Four-rank reload cases with aligned checkpoint and batch Known EOS occurrences, both arms Request preemptions, both arms

50 4/4 36/36 9,472 10/10 0 0

50 4/4 36/36 9,472 10/10 0 0

ledger task (opening balance 137, receipt 58, dispatch 29) with repeated background text. Its 12 tokenized inputs vary length and early document identity; nine inputs enter the measured sample. They are input variants of one synthetic task, not 12 distinct tasks. Requests use temperature 0, seed 42, and ignore_eos=True; task-answer quality is not evaluated. TTFT measures request start to receipt of the first nonempty token-ID event; total latency ends after the stream’s DONE marker. Decode rate counts tokens after the first event over the first-to-last nonempty-event interval, allowing multiple tokens per event. CPU metadata observers remain enabled; GPU tensor probes, poisoned destinations, artificial save delays, initialization and interrequest pauses are excluded. The GPUs are dedicated to the experiment while other production workloads continue on the shared host. Trials run in fixed ascending order of trial index and prompt length; only engine order is reversed. This balances engine order across two runs but does not randomize request history or eliminate shared-host interference. The performance runs use four NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. Both arms use the Marlin MoE backend with custom allreduce disabled. NCCL_P2P_DISABLE is set to 1 and OMP_NUM_THREADS to 2; expandable_segments is False. FlashInfer autotuning is disabled while the separately fixed FLA profile remains in force. The candidate’s local LMCache server uses LRU, separate object groups, l1-size-gb=16, and disabled lazy L1 allocation. These settings bound the performance result; no inter-host cache transfer or optimized communication comparison is measured.

Each latency is the median of three measured trials. Ratios are medians of per-input OFF-cold/ON-reload ratios, not ratios of the displayed medians. Across the six run/length groups, TTFT ratios range from 1.85–2.80 and total-latency ratios from 1.019–1.076. ON-cold/OFF-cold TTFT ratios range from 1.01–1.04; this is an end-to-end enabled/disabled comparison under the shared modified controls. It does not isolate transfer latency, observer cost, or the cost of all changes relative to production. Raw perrequest values, ranges and event-level decode rates are retained in the evidence tables. Across these same six groups, OFF-cold total-latency medians span 5.570–6.104 s and ON-reload medians span 5.369–5.745 s. Post-first-event decode-rate medians span 23.38–24.80 and 23.08–24.76 tokens/s, respectively. These ranges describe the observed time scale across inputs and runs; they are not uncertainty intervals. Attempted path names do not determine actual paths. Among 18 measured OFF immediate repeats, 12 recompute and 6 reuse a local prefix; among 18 ON immediate repeats, 18 reload from CPU and 0 use local-only reuse. The OFF local reuse occurs only at N=5376 and credits 1792 tokens; its N=3583 and N=3584 repeats recompute. At N=5376 the ON CPU path instead restores 3584 tokens, so those repeat conditions do not perform equal suffix work. Consequently, a complete five-path local-hit comparison is unavailable. This is an observed limitation of the evaluated configuration, not evidence that hybrid models cannot support local reuse. Each of the 18 measured explicit CPU reloads has positive fourrank H2D and an independently verified first-batch prefix B=p: 1792, 1792 and 3584 tokens for the three lengths, leaving 1791, 1792 and 1792 prompt tokens to recompute. Classification retains this actual suffix work. We report

6

a Time to first token (ms)

b Full request time (s) 7 6

600

5 4

400

3 2

200

1 0

0 3,583 R1

3,584 R1

5,376 R1

3,583 R2

3,584 R2

5,376 R2

3,583 R1

3,584 R1

Prompt length N / run OFF cold

5,376 R1

3,583 R2

3,584 R2

5,376 R2

Prompt length N / run ON cold

CPU reload

Figure 3: Serial latency measurements. Hollow points show three measured requests per run, length, and condition; filled markers show medians, and thin bars show observed min–max ranges, not confidence intervals. R1 runs OFF then ON; R2 reverses engine order. Both panels use the same 54 requests and zero-based axes. OFF is the modified no-connector control. The cold conditions have similar TTFT; total time includes the full 128-token generation. Table 7: Serial latency under matched controls Run

N

OFF cold TTFT (ms)

ON cold TTFT (ms)

ON reload TTFT (ms)

TTFT ratio

Total ratio

1 1 1 2 2 2

3,583 3,584 5,376 3,583 3,584 5,376

447.3 447.6 670.9 447.3 447.2 671.2

451.9 458.8 686.7 455.0 462.6 691.8

239.4 240.4 239.7 237.9 241.4 241.2

1.87 1.86 2.80 1.88 1.85 2.78

1.030 1.037 1.076 1.019 1.031 1.053

no confidence interval from two engine pairs, sustainedconcurrency result, or capacity gain beyond GPU memory.

speculative KV state with verified state even when generated tokens agree. These works reinforce why finite token equality should not be promoted to complete state or future-output identity. Our experiment fixes one observed profile and tests finite continuations; it does not implement a general online verification system [2], [5]. LMCache supplies the cache infrastructure used in this study. Its broader cache management and transfer mechanisms are existing platform contributions. The local cumulative patch spans representation, completion, scheduling, and common computation changes; eleven changed files do not constitute eleven independently novel mechanisms. Only the final lookup intervention is isolated by the otherwise controlled before/after matrix. The preceding integration changes were not subjected to a full factorial ablation [6]. External validity is limited by one model revision, one four-rank setup, synthetic prompt templates, finite continuations, serial requests, and a fixed profile selected from an observed baseline. The original matrix uses one template and 64-token outputs; the extension adds three templates and 256-token outputs, with only two freshcontainer repetitions. We have not measured all logits, arbitrary future continuation, cross-TP invariance, concurrent preemption, other models, or default-production equivalence. The complete modified control is essential to interpreting the result.

6 Related work and limitations Hybrid-state reuse is an established problem. Marconi explains why recurrent states updated in place require checkpoints aligned with reusable attention prefixes and studies cache admission and eviction. The Sparse Prefix Caching preprint studies sparse checkpoint materialization and suffix recomputation for hybrid and recurrent serving. Our strict-prefix lookup repair applies this existing requirement to a connector/scheduler disagreement. We introduce neither a new eviction policy nor the principle of replaying a suffix from a checkpoint [1], [3]. An upstream vLLM checkpoint RFC separately discusses prefix identity, non-KV logical state, and immutable payload versus logical equivalence. It is a design proposal, currently closed as not planned, rather than an implemented standard. Its prior discussion limits a novelty claim about a general recovery interface; the contribution here is evidence from a running integration [4]. Numerical-control work is also directly relevant. He distinguishes repeatability from batch invariance. The LLM-42 preprint uses verified speculation and replaces

7

References

The serial study measures a CPU-reload latency benefit against the shared modified cold control. Costs of fixed configuration, ordering changes and completion barriers relative to the original production configuration remain unmeasured. Immediate repeats do not consistently provide local-only reuse, limiting a complete five-path comparison. Two engine pairs on one shared host do not establish concurrent SLOs, broad reliability, or capacity beyond GPU memory. No public artifact release or production deployment is claimed.

[1] Rui Pan et al. Marconi: Prefix Caching for the Era of Hybrid LLMs. MLSys, 2025. [2] Horace He and Thinking Machines Lab. Defeating Nondeterminism in LLM Inference. Author technical report, 2025. [3] Mikhail Shirokikh and Sergey Nikolenko. Sparse Prefix Caching for Hybrid and Recurrent LLM Serving. arXiv:2605.05219v1, preprint, 2026.

7 Conclusion The evaluated GLM integration restored a complete checkpoint while crediting the scheduler with a shorter prefix. Actual transfer metadata and batch records exposed that mismatch; restricting lookup to a strict prefix repaired both failing exact-boundary cases under a controlled numerical configuration. The repaired nine-case matrix passes all 36 finite continuation comparisons, and a separate instrumented run passes the specified byte, nonzero-tail, and delayed-save checks. The subsequent content extension also passes two fresh-container comparisons with 256-token continuations. The engineering result is a validated recovery path within these conditions, with a measured serial CPU-reload latency benefit under matched controls; representative concurrency, originalproduction overhead and overflow capacity remain open.

[4] vLLM checkpoint RFC, issue #40533. Engineering design proposal, 2026; closed as not planned when reviewed on 14 September 2026. [5] Raja Gond et al. LLM-42: Enabling Determinism in LLM Inference with Verified Speculation. arXiv:2601.17768v2, preprint, 2026. [6] Yuhan Liu et al. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference. arXiv:2510.09665v1, preprint, 2025.

Acknowledgements

[7] PVC (Research Infrastructure), UNSW Sydney. Katana. UNSW, Sydney, 2010. doi:10.26190/669xa286.

The author acknowledges Research Technology Services, UNSW Sydney, for providing GPU resources on the Katana computational cluster [7] for model quantization.

[8] Z.ai. GLM-5.3-Flash. Model card, accessed 14 September 2026.

AI assistance. AI tools assisted with the research and preparation of this manuscript. The author takes responsibility for the methods, results, and final text.

[9] Red Hat AI. GLM-5.3-Flash-NVFP4. Model card, accessed 14 September 2026.

8

Record · ID 919349 · SHA-256 016063efac9b3be9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.