Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys Georgios Politis
Setloop.io
arXiv:2609.04382v1 [cs.CR] 3 Sep 2026
Evangelos Pappas [email protected]
Abstract We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested. The Trusted Local Node (TLN) sends protected activations to the Untrusted Cloud Node (UCN), the UCN returns its output, and TLN, holding the private loss, returns the output gradient. The frame the UCN receives mixes real rows with decoys, and the loss ignores the decoys. Their gradients are exactly zero, so the pattern of zeros reveals which rows were real. We measure it with a protocol fixed in advance: a leak injected at known strength to prove the instrument can see one, a shuffled-label control to prove it does not report absent leaks, and a threshold set before the runs. Across nine seeds, the zeros identified the real rows on every frame, 4,096 of 4,096 per run. An attack on the frame contents recovered about one extra token per hundred over a constant-guess baseline (+0.65 to +1.50 percentage points); the shuffled controls recovered nothing. A second set of runs repeated this on a configuration that keeps model quality within budget, so the finding is not confined to a setting nobody would deploy. On both datasets, every such run passed the forward-channel privacy check and the quality check, yet failed that same check once the returned gradient was included. Clipping and noising each row of the gradient closed the leak for about 0.01 nats of held-out cross-entropy. The system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.
1
Introduction
Split learning [31] lets a data owner rent cloud compute for training without sending raw examples: the trusted side sends activations and, because it alone holds the private loss, returns the output gradients the untrusted side needs to train. The privacy claim of such a system is that the cloud cannot read the training data. The claim rests on an evaluation, and the evaluation rests on an instrument. The instrument examined here failed silently. The original evaluation of this two-node split-LLM system instrumented the forward wire, the activations sent to the cloud, and passed its privacy gate. The backward wire, which carries the output gradient the trusted side returns to the cloud, was never declared a privacy surface. In this implementation, excluding the decoy rows from the private loss makes their returned gradients identically zero; that zero-support construction is an implementation and system-design defect. The false assurance has a distinct, more general evaluation cause: the observable gradient channel was outside the declared adversary view, so the gate never tested it. This paper is therefore a systems-security case study with a calibrated protocol applied to one system, not a general methodology paper. Its contribution is a verified diagnosis of that case and a channel-explicit evaluation, positioned against three recent split-LLM evaluations, attack and defence alike (Section 7). Separating the real rows from the decoys matters because the decoys exist precisely to hide which rows carry the private loss. Exact gradient support lets the cloud discard all 48 decoys and isolate the 32 loss-bearing rows before any content attack, collapsing the anonymity set that padding was meant to create. 1
This is a structural metadata disclosure, not text reconstruction; the content evidence is separately bounded to the implemented frequent-token probe. Section 2 defines the system and its declared threat surface. Section 3 lays out the evaluation protocol, and Section 4 calibrates the instrument against injected leaks. Section 5 reports the structural disclosure, its replication, and the mitigation runs; Section 6 reports how the signal converts to a token advantage with depth, width, and budget. Section 7 audits three external evaluations and places the work against the DP-auditing and split-LLM literature. Section 8 records scope, limitations, and the verification trail.
2
The system and the threat model
Two cooperating nodes train one LLM. We refer to them throughout by their roles: the Trusted Local Node (TLN) holds the data and the embedding head, and the Untrusted Cloud Node (UCN) holds a middle stack of transformer layers. For each training frame, TLN sends a protected latent (the forward wire); UCN trains its stack and returns its output; TLN, which alone holds the private loss, scores that loss and sends the output gradient back to UCN (the backward wire). Both tensors are held by the untrusted node in the clear. The defence on the wire is a latent-space bottleneck at width D=64, a per-request rotation and permutation, decoy rows mixed among the real ones, and boundary clipping and noise. Each 80-row frame carries 32 real rows and 48 decoys. The implementation and the artefacts call these decoys chaff, after the radar countermeasure, and we use the two words interchangeably. The original evaluation declared the forward wire as the privacy surface, instrumented it, and passed its gate. The backward wire was outside the declared adversary view. Table 1 sets out the evaluated configuration together with the replication and mitigation runs. field
evaluated setup
model and runtime
Qwen3-0.6B; Transformers 5.13.0 runtime; exact weights SHA-256 f47f71177f32bcd101b7573ec9171e6a57f4f4d3 1148d38e382306f42996874b Main configuration and replication runs: split after layer 14, resume at 26 (11 delegated layers). Mitigation runs: split 21/26 (4 delegated layers). WikiText-2; corpus SHA-256 78b6bfb90cfd718f0c27d42b1fd2231b139d1dda75d7d796e6a603b2e5cd7efe. The runner tokenizes the flat corpus into disjoint, sequential fixed-width blocks, then slices train followed by evaluation; this is not a document-level split. 32 real + 48 decoy = 80 rows per frame; D=64. Forward and returned activations: per-row C=1.0 and Gaussian σ=0.35C. Main configuration and replication runs, outbound gradient: open. Mitigation outbound gradient: C=0.01 and Gaussian σ=0.35C. coordinate plus invariants, invariant only, and invariant graph; 20 epochs; batch 512; 3 restarts.
topology corpus and implemented split frame and latent boundary protection known-plaintext latent probe captured frames per run seed groups decision thresholds inference
Main configuration and replication runs: 4,096 = 2,048 probe-train + 2,048 evaluation. Mitigation runs: 1,024 = 512 probe-train + 512 evaluation. 42–47 exploratory; 48–50 replication; 51–53 mitigation. Privacy: gate statistic > +1.0 pp fails. Utility: held-out cross-entropy increase ≤ 0.35 nats passes. Inferential unit: one independent training seed. Within-run uncertainty is frame-clustered.
Table 1: Experimental setup for the main configuration and for the replication and mitigation runs. Full fingerprints identify the exact model weights and corpus used by the evaluated runner.
The threat model is a fully compromised remote node: UCN can read every message, state update, and cross-step observation on its own side. Prior split-learning work demonstrates passive reconstruction and label leakage as well as active backward-signal manipulation [9, 18, 24]. Accordingly, the declared adversary view has multiple channels, and a privacy claim is only as strong as the enumeration and testing of those channels. Figure 1 depicts the protocol’s three hops, Figure 2 expands one frame into its per-step anatomy, and Table 2 records the trust split at the evaluated operating point.
2
1. forward wire (activations) clipped and noised · gate passed TLN (trusted)
data · head · tail holds the loss
2. return wire (remote output) clipped and noised · not attacked
UCN (untrusted, cloud)
middle layers 3. backward wire (output gradient) unclipped, unnoised · leaks
Figure 1: The three hops of one training step. Only TLN can compute the gradient, because only TLN holds the loss, so hop 3 travels left to right just as hop 1 does: “backward” names the pass the tensor belongs to, not the direction it moves. Hops 1 and 2 are clipped and noised, and the original evaluation attacked hop 1 and passed its gate. Hop 3 crossed raw and carried the leak. The zero-support signal is an implementation and system-design defect; the false pass is an uninstrumented-channel evaluation failure.
3
TLN (trusted)
UCN (compromised)
S1. Session setup. Corpus → tokenizer → 32-row blocks; gauge masters drawn from SHA-256 counter-mode KDF (128bit).
S2. Fresh cloud surrogate + AdamW per session (∼100 params, gauge-equivariant Gram kernel).
SEALED: corpus, encoder/decoder, masters, labels
1. Frozen prefix. h = prefix(x), H=1024. NEVER crosses
2. Private encoder. z = E(h), D=64. NEVER crosses
3. Fresh gauges. z̃ = Π diag(s) z R: rotation R, permutation Π, scale s (fresh per request, KDF-derived). OBFUSCATED: R, Π, s
4. Decoy mix. +48 recycled real rows → 80 gauged rows. OBFUSCATED: decoys
5. Perturb. clip ∥·∥ ≤ 1.0 + Gaussian σ=0.35 (both directions).
6. forward wire: 80×64 frame, TLS 1.3
PROTECTED: clip + noise
HELD IN CLEAR
7. Surrogate forward on the 80 gauged rows → output. 8. return wire: remote output (clipped, noised)
compromised compute
9. Private loss → output gradient g. Decoy rows get identically zero gradient (loss truncates before decoys). MECHANISM BORN: zeros = partition
10. backward wire: g UNCLIPPED, UNNOISED (pre-fix)
⋆ LEAK POINT
12a. TLN step. Encoder/decoder update (minimax vs. probes).
11. Adversary’s read. zero-support ⇒ exact real/decoy split (1,024/1,024 frames) ⇒ drop decoys ⇒ joint view (frame ∥ gradient). 12b. UCN step. AdamW on the surrogate.
Figure 2: Anatomy of one training cycle, per frame. SEALED never crosses the boundary; OBFUSCATED/PROTECTED are the defence’s active layers; LEAK marks the unprotected backward wire and the partition mechanism it discloses. The forward probe passes at step 6 while steps 10 and 11 leak: the paper’s finding in one diagram.
4
TLN (trusted) keeps
UCN (compromised) sees/holds
Private corpus; token IDs; plaintext I/O
Gauged D=64 frame rows only (80 rows/frame: 32 real + 48 decoy) A ∼100-parameter gauge-equivariant surrogate it trains itself Unclipped, unnoised D-width output gradients (the omitted channel) Protocol metadata: frame sizes, timing, session shape digests TLS 1.3 endpoints and ciphertext (pinned CA)
Frozen LM prefix, embedding, LM tail Private encoder H→D=64 and decoder D→H weights Per-request gauge masters (rotation, permutation, scale; SHA-256 counter-mode, 128-bit) Honest evaluation labels; session keys; optimiser state for trusted parts
Never crosses: H-width activations, canonical coordinates, token order, token scale, plaintext tokens. Table 2: Trust split at the evaluated operating point. The output gradient crosses the backward wire unclipped and unnoised.
Terminology. Four project terms appear throughout without prior-art homes, so we define them once here (no citations; they are constructs of this system, not borrowed results). A seed is the random seed initialising one training run and, through the KDF, every per-request gauge draw — the unit of independent replication. A cell is one complete experiment — a defence configuration, a seed, and a training budget, run end-to-end and then attacked and scored. A battery is a predeclared set of attacks scored as one sweep: the frozen nine-arm probe family (three model classes × three restarts) or the compromise-fraction sweep. A surrogate is the small gauge-equivariant module UCN trains in place of the real middle layers (∼100–161 parameters) — a stand-in that computes on gauged frames without ever learning the gauges. Likewise: the gate is the predeclared decision threshold, the floor is the reading of a matched no-attack control, an arm is one scored attacker instance, chaff is the recycled real decoy rows, and a gauge is one fresh per-request randomisation (rotation, permutation, or scale).
3 Evaluation protocol: declared channels, calibrated metrics, and the gate statistic The gap this paper addresses is not discovery of gradient leakage or a bidirectional attack surface: both are established in gradient-inversion and split-learning work [36, 7, 6]. The gap is a concrete evaluation returning a pass without testing whether its metrics can detect a known leak on every declared channel. We instantiate three established audit disciplines for that setting. Declare every channel. The evaluation must enumerate all channels the adversary observes (forward, backward, membership, timing) before it measures any of them. A channel absent from the declared view is exempt from the gate by construction; the leak found here lived exactly in that exemption. Table 3 records the enumeration and its current status. Its artefact sources and verification procedures are indexed in Appendix A.
5
family
applicable metrics
positive control
status
forward only
injected leak on the forward wire; codeword injection injected leak on the backward wire
measured
joint forward gradient
token top1, rare token top1, token cross entropy token top1, rare token top1, token cross entropy token top1
measured
accumulated history
token top1
scaled joint concatenation (the exploratory frequent-token gradient result) none
stateful remote state
token top1
none
timing metadata
—
none
active perturbation
token cross entropy
amplitude sweep
membership property
—
none
response side
—
none
gradient only
measured
unmeasured (no positive control) unmeasured (no positive control) unmeasured (no applicable metric) constructible, unexecuted unmeasured (no applicable metric) unmeasured (no applicable metric)
Table 3: The declared channel families and their status. “Unmeasured” means the family has no applicable metric or no positive control, so it cannot support a primary claim.
Calibrate every metric against a known leak. A metric certifies the absence of a leak only if it detects a leak injected on purpose. Section 4 measures each metric’s detection threshold with a controlled dose– response sweep. A metric that has not been calibrated cannot stand between a privacy claim and a passing verdict. Planted canaries and attack-based privacy audits provide the methodological precedent [4, 15, 22, 29, 26]; our adaptation makes the dose and decision channel- and metric-specific. Define the gate statistic and paired effects. Every attack result in this paper is a top-1 token accuracy: the share of evaluation tokens the attacker’s probe names correctly. We report it not as a raw accuracy but as the gap, in percentage points (pp), between the probe and a constant baseline that always guesses the single most frequent token in the evaluation set. That baseline sits near 5 to 6% for this model and corpus, so an effect of +1 pp means the attacker recovers roughly one extra token per hundred, about a sixth more than guessing alone would give. The gate is set at +1.0 pp, and the effects measured here fall between about +0.5 and +2.3 pp. Let J be the nine predeclared probe arms (three probe variants, each scored at three restarts; Table 1), Bonf U0.95 (b pj ) the Bonferroni-adjusted Wilson upper-95 accuracy for arm j, and pbconst the point accuracy of the constant baseline, which always predicts the most frequent evaluation token. The historical gate statistic is Bonf G = max U0.95 (b pj ) − pbconst . j∈J
(1)
Thus the constant baseline is subtracted after the Wilson bound is formed; G > +1.0 pp fails the gate. For frame f , define the paired difference ∆f (q) = accf (q) − accf (constant baseline).
(2)
For a real arm r and its shuffled-label negative control s, respectively, F
F
1 X A= ∆f (r), F
1 X N= ∆f (s). F
f =1
f =1
6
(3)
A is the real-arm paired effect; N is the same paired effect for the shuffled-label negative control. N is a false-positive check and is not subtracted inside G. Within-run uncertainty bootstraps frames; reported across-seed contrasts, including A − N , use a hierarchical bootstrap over seeds (outer) and frames (inner). Three verdict terms recur throughout, and we fix them here. A result is detected when its lower bootstrap bound is positive or when exact support evidence establishes it; an arm is at floor when that bound includes zero, so the arm is indistinguishable from its constant baseline; and an arm breaks the gate when, and only when, G > +1.0 pp under the Bonferroni–Wilson rule of Equation (1). “Detected” and “breaks the gate” are not synonyms: a paired effect can be detected well below the gate, and the gate can break on an upper bound whose point effect is smaller. We also call one (configuration, seed) run a cell, and one capture run’s serialised frames plus its metadata manifest a bundle. z-style row-independent statistics remain only for comparability with previously reported results of this system. Calibrated reference distributions and explicit operating points are established in membershipinference evaluation [3, 28]; the shuffled-label negative control is their channel-specific counterpart here.
4
Instrument calibration
Calibration is per metric: thresholds are set only after each metric’s detection curve is measured. The sweep injects a known token leak at controlled dose (coverage × amplitude) and reads where each metric detects it.
4.1
The four metrics disagree
Figure 3 and Table 4 show the dose–response curves. token top1 has a sharp onset between coverage 0.04 and 0.06, steepening through 0.10. rare token top1 is the most sensitive, first responding at coverage 0.04 and reaching full recovery at coverage 1.0. token cross entropy is dose-insensitive until the injected leak dominates. membership auc is insensitive over the tested low- and moderate-dose region (AUC−0.5 spans 0.058–0.159 across the full sweep), because the injected leak is a token-identity signal, not a membership signal; the metric was subsequently falsified as a channel and retained as a probe-generalisation diagnostic only. The curves disagree, consistent with broader evidence that reconstruction metrics need not agree on privacy risk [30], so no single threshold fits all metrics. The thresholds (Table 5) are set per metric, and they are budget- and frame-invariant along the remaining sweep axes. Calibration changed how the decision is interpreted, not the gate formula: three cells later shown to be degenerate (arm identical to the constant baseline row-for-row) had pinned the statistical floor, and the sweep measures that false-negative region while preserving the historical +1.0 pp Bonferroni–Wilson rule. The frequent-token effect is detected on every seed (A from +0.69 to +1.19 pp, |N | ≤ 0.08 pp); explicit gate breaks are reported only where the Bonferroni–Wilson column exceeds +1.0 pp (Section 5).
7
100
token_top1 rare_token_top1
75 50 25 0
+1.0 pp gate
10−2
10−1
injected coverage (fraction of rows)
100
metric value (nats / AUC-0.5)
recovery advantage (pp)
dose-response per metric: the four curves disagree token_cross_entropy membership_auc (AUC-0.5)
6 4 2 0
10−1 10−2 injected coverage (fraction of rows)
100
Figure 3: Dose–response per metric on the coordinate-mode sweep (amplitude 1.0). The four emitting metrics disagree: rare token top1 is most sensitive; token cross entropy barely moves until the leak dominates; membership auc does not detect the low-dose token signal because the injection is not a membership signal. A single calibration run across 19 doses.
metric token top1 (pp) rare token top1 (pp) token cross entropy membership auc (AUC−0.5)
cov 0.02
0.04
0.06
0.10
0.30
1.00
0.35 0.08 1.114 0.093
0.35 0.32 1.114 0.093
0.98 0.67 1.114 0.095
4.19 1.97 1.115 0.095
22.51 14.18 1.130 0.096
94.88 100.0 7.025 —
Table 4: Coordinate-mode coverage sweep at amplitude 1.0. The calibration reference is the zero-dose source cell (coverage 0, same bundle family) and the constant baseline (the most common evaluation token); the calibration sweep contains no shuffled-label arm, because those negative controls are introduced only at scoring time. The top-1 rows show point effects, while the gate statistic uses their Bonferroni–Wilson upper-95 value; thus the gate first breaks at coverage 0.06 although the displayed point effect is 0.98 pp. Cross-entropy is a proper scoring rule; the values shown are the effect over the constant baseline, and at scoring time CE is read arm minus its shuffled-label negative control, never raw (Table 5). membership auc values are flat. “—” marks the amplitude-axis section, not part of this row.
metric
threshold set in advance
basis
token top1 rare token top1 token cross entropy
gate: Bonferroni upper-95 excess > +1.0 pp paired effect +1.0 pp; report raw recovery none set; read arm minus shuffled-label negative control (never raw) none (diagnostic only)
first displayed gate break at cov. 0.06 earliest subthreshold movement at cov. 0.04 a confidently-wrong negative control inverts raw CE; no numerical gate falsified as a channel; constant baseline vacuous
membership auc
Table 5: Per-metric thresholds, set in advance. A metric without a measured detection curve may not appear in a primary claim.
4.2
The representation-matched positive control
A floor reading is interpretable only if the same probe succeeds on a representation-matched undefended control. The isolation-audit bundles do not provide that control because their naked boundary is a different representation and the probe remained near floor there. For the split-14 claims, the positive control therefore uses the split-14 naked capture and a probe built to detect it: a four-layer, eight-head Transformer encoder trained 50 epochs to map each released latent row back to its token . Appendix A indexes the associated artefacts. 8
boundary naked (D=1024, no defence) defended (D=64, noise, decoys) defended (D=1024, dimension-matched)
Wilson upper-95 acc.
constant baseline
excess
note
28.54% 5.56% 5.56%
4.35% 5.27% 5.29%
+24.19 pp +0.29 pp +0.27 pp
breaks at floor at floor
Table 6: One probe architecture, one protocol, three boundaries. The probe that inverts the naked boundary at over six times its baseline reads nothing through the defence, at the operating width and at matched width, so the defended floor readings are not attacker blindness. Best accuracy is selected on the evaluation set over 50 epochs (a selection reading); the accuracy column reports its Wilson upper-95 bound, so each row’s excess is the displayed accuracy minus the baseline. The defended probes’ accuracy declines after its early epochs, consistent with overfitting noise rather than extracted signal. The committed three-restart rerun of the D=64 defended probe reads at the same floor (+0.37 pp); the other three committed deep-probe runs are the rows above.
The naked break selects the best epoch on the evaluation set, so +24.19 pp is a sensitivity demonstration, not a primary leak estimate; and the defended frame carries 80 rows against the naked frame’s 32, matching the study’s evaluation convention. What the contrast establishes is narrow and load-bearing: an attacker architecture proven sensitive to the representation reads nothing through the defence, while an attacker family that could not read even the naked boundary could never have shown it. Boundary-specificity is measured: re-run on the packaged isolation-audit bundle, the same probe reads the naked boundary at floor as well (+0.66 pp, its shuffled-label negative control at +0.50). This is exactly why a positive control must be representation-matched to the cell under test, and why the split-14 capture above is the control for the split-14 claims.
5
Results: structural gradient leakage and bounded content inference
Table 7 summarises the main configuration and the result categories developed in this section; availability and verification details are indexed in Appendix A.
9
result category
result
interpretation
forward wire backward wire
gate passed leaks structurally
frequent-token content effect nine audit cells
pooled over the nine verified seeds: +0.92 pp, 95% interval [0.74, 1.09] effect detected at 12 and 11 delegated layers, at floor at 8 and 6; detected at D=64 and D=96, at floor at D=128; 2.5× more exposure does not amplify token top1 onset ∼0.04–0.06; rare token top1 at 0.04; CE flat until dominant; membership auc flat and dose-insensitive all nine cells ≤ 10−6 pp
the original instrument’s only channel 4,096/4,096 frames; row agreement 1.000, on each of nine seeds negative control at floor on every seed; per-seed values in Table 8 one run each (seed 42); partition exact at every depth
dose–response calibration
internal scorer reimplementation mitigation runs
joint view breaks the gate on 6/6 defended cells across two datasets; gradient clipping and Gaussian noise suppress the scored probes for a held-out cost of ≈ 0.01 nats
per-metric thresholds set in advance
internal check (Appendix A)
under the fixed protocol; four-layer topology
Table 7: Result categories and where each is established. The final row refers to the four-layer mitigation topology; every row above it refers to the primary configuration.
5.1
Mechanism: the gradient says which rows are decoys
A decoy row’s gradient is identically zero, because the trusted-side loss truncates before the decoys. With the gradient unprotected, the zero-support pattern of each returned gradient frame exactly repeats the split between real rows and decoys. On every one of 4,096 frames, on every seed, the match is exact: 4,096/4,096 frames; row-level agreement 1.000. The corresponding per-seed artefacts are indexed in Appendix A. The result has three claim levels. (i) Structural metadata leakage. Gradient support discloses exactly which rows are real: the cloud can separate every loss-bearing row from every decoy row in every evaluated frame. This is the disclosure the backward wire is uniquely responsible for. (ii) Content inference. Conditioned on that structural signal, the pre-set frequent-token probe has a modest paired effect. The gradient’s contribution here is small: comparing the gradient and forward arms of Table 8 seed by seed, the gradient’s margin over the oracle-partitioned forward frame is at most +0.18 pp on five of the six exploratory seeds (+0.0011 to +0.0395 on four of them), and reaches +0.64 pp only on seed 42. The content an attacker reads therefore sits largely in the forward frame, and what the backward wire supplies is the partition that makes the frame readable, which is claim (i), not additional content. This probe result is content inference; it is not evidence of transcript or sequence reconstruction. (iii) Not established. The experiments do not establish rare-token recovery, sequence reconstruction, or held-out-text reconstruction. Those claims require different emitters and controls and remain outside the measured result.
10
The cycle in which this happens is Figure 2, steps 9 to 11. Table 8 reports the exploratory and replication runs. seed
gradient arm (pp)
forward arm (pp)
negative control (pp)
frames exact
verdict
42 43 44 45 46 47
+1.1923 +0.6893 +0.6929 +0.8435 +0.9361 +0.7705
+0.5516 +0.6821 +0.6918 +0.8040 +0.7578 +0.7618
−0.0069 −0.0231 +0.0479 −0.0707 −0.0285 −0.0308
4,096/4,096 4,096/4,096 4,096/4,096 4,096/4,096 4,096/4,096 4,096/4,096
detected detected detected detected detected detected
mean
+0.8541
+0.7082
| · | ≤ 0.08
4,096/4,096
—
48 (fixed protocol) 49 (fixed protocol) 50 (fixed protocol)
+0.6456 +1.5008 +0.9929
+0.5023 +0.8721 +0.6326
−0.0571 +0.0232 +0.0650
4,096/4,096 4,096/4,096 4,096/4,096
detected detected detected
Table 8: Exploratory seeds 42–47 and the replication runs (seeds 48–50). Statistic: the paired effect of each arm over its constant baseline (the most frequent evaluation token), clustered by frame. The negative-control column reports the label-shuffled twin of the gradient arm, scored through the same statistic. Inferential unit: one training seed; each row one independent run, its interval a cluster bootstrap over that run’s frames. All nine committed seeds are shown. The forward arm is built on the oracle real/decoy split and therefore does not by itself describe an achievable attack. It is reported here because in this system the oracle is redundant: the gradient reproduces the partition exactly (Section 5.1), so the forward arm is what an attacker reads after the backward wire has supplied the partition. No label-shuffled control was run for the forward arm, so that column is uncontrolled; the negative-control column bounds the gradient arm only. Seed 42 was served by a cloud container built before the code was packaged for release; seeds 43–47 ran the released package on both nodes, and seeds 48–50 are captures made under the fixed protocol.
5.2
What this does, and does not, establish
The exact partition and the modest frequent-token paired effect are detected under a shuffled-label negativecontrol protocol the audited evaluations do not run. It establishes that a gate that never declares the backward channel cannot see the partition signal. The originally published gate would also have passed this cell under its own threshold. The structural disclosure was invisible to it twice over: the channel was never scored, and the floor the threshold was calibrated against was degenerate (Section 4).
5.3
Does the leak survive a configuration worth deploying?
The main configuration fails its own utility gate: it protects the data but costs too much model quality to be worth running. That leaves an objection. If the leak only appears in a configuration nobody would deploy, it may not matter. This section answers that objection with a second set of runs on a configuration that does pass both gates, so the leak cannot be dismissed as an artefact of an unusable setting. These runs use a shallower split, delegating four transformer layers instead of eleven. They were run on three fresh seeds (51 to 53) from the released code, and scored against the thresholds set in advance, so nothing about them was tuned after the fact. Table 9 reports every run in the mitigation set.
11
cell (seed)
utility ∆
forward gate
joint arm (gate)
joint arm (paired effect)
support leak
defended, gradient open (s51) defended, gradient open (s52) defended, gradient open (s53)
+0.054
+0.522
+1.954
+1.609
1,024/1,024
+0.063
+0.574
+2.180
+2.333
1,024/1,024
+0.071
+0.412
+1.765
+1.567
1,024/1,024
defended, gradient clipped and noised (s51) defended, gradient clipped and noised (s52) defended, gradient clipped and noised (s53)
+0.063
+0.406
+0.480
−0.204
0/1,024
+0.078
+0.421
+0.476
−0.017
0/1,024
+0.083
+0.412
+0.481
−0.013
0/1,024
public corpus, gradient open (s51) public corpus, gradient open (s52) public corpus, gradient open (s53)
+0.092
+0.385
+1.632
+0.986
1,024/1,024
+0.080
+0.767
+1.607
+0.942
1,024/1,024
+0.083
+0.483
+1.435
+0.750
1,024/1,024
+0.00/ + 0.12/ + 0.43†
+9.5 to +10.7
+33.9 to +34.0
≈ +51
—
naked control (s51–s53)
Table 9: The mitigation runs. The gate columns report the gate statistic: Bonferroni-upper-95 excess over the constant baseline. The paired column is the cluster-bootstrapped paired effect of the best arm against its constant baseline. “Joint arm” is the forward frame and the output gradient of the same training step, concatenated: the view the compromised node actually holds, and the one the original gate never scored. This set comprises fourteen runs; the twelve scored cells are shown. The two omitted are a configuration smoke test that produced no scored arms and the packaged seed-42 rerun of the main configuration, reported separately in this section. † The naked control on seed 53 fails the utility gate (∆ loss 0.432 against the 0.35 bound); it is retained because it is a deliberately undefended sensitivity control, not a candidate configuration. In every one of these runs the per-row scale-and-sign gauge was disabled, as it is for the rotation-invariant cloud architectures (Appendix B); the joint-view results below are conditional on that setting.
(a) The joint-view gate result is not confined to a utility-failing configuration. On all six defended, gradient-open cells (three seeds on each of two datasets) the run passes the forward privacy gate and the utility gate, and the joint view still breaks the +1.0 pp gate set in advance, while the gradient on its own stays at the level of guessing. (b) The mitigation removes the zero-support partition signal. With per-row gradient clipping and Gaussian noise, every scored arm sits at floor and the structural partition disclosure is gone: 0/1,024 frames, agreement 0.400 at the 32/80 base rate. Utility is intact, at a held-out cost of ≈ 0.01 nats. (c) The naked control restores the signal. The same topology with the defence off reads +34 pp on the joint view, so the defended cells’ near-floor readings reflect the defence rather than a blind instrument. On the forward gate the defended cells sit at floor while the naked cells exceed it by roughly ten pp, separating the two configurations by more than an order of magnitude. Cross-Gram cancellation does not explain the joint-view result. The forward activation matrix X and its returned gradient G share the per-request rotation, so their cross-Gram product XGT cancels it. The pre-set probe family did not compute this cross-tensor feature. We implemented it (cosine and scale-keeping forms, with rotation-invariance self-tests) and scored it on every captured cell: it reads at floor everywhere, defended (gate statistic +0.31 to +0.86, paired effect −0.29 to +0.46) and naked alike (paired effect −0.78 to −0.22). Under the two implemented emitters, the joint-view effect therefore runs through the concatenated gradient block’s own content, not through the rotation-cancelling cross-term; other cross-tensor attacks remain open. The effect exists on a second corpus. The main results are on the private WikiText-2 slice. The
12
alternative public WikiText-2 corpus carries the same topology on seeds 51–53 (Table 9), and on two of the three the gradient alone breaks the gate as well (+1.175 and +1.219). One additional corpus is a limited robustness check: the effect is observed on both tested datasets, at different magnitudes, which does not establish that the effect is dataset-independent. Packaged seed-42 rerun. The seed-42 result in Table 8 (+1.1923, originally served by a pre-existing container) was re-run from packaged code in a clean container in the primary configuration: +1.147 (upper95), paired effect +1.017, support exact on 4,096/4,096, and the negative control at floor. The result is not confined to that container. Pooling the nine seeds that were run from packaged code (the six exploratory and the three replication runs), a seed-level random-effects estimate puts the frequent-token gradient-arm paired effect at +0.92 pp, 95% interval [0.74, 1.09] (τ = 0.23, I 2 = 77%). The interval reaches the +1.0 pp gate, so the pooled content effect is not separable from the gate threshold at this sample size. Three further units from an earlier, unpackaged tree are retained in the committed estimate for continuity with earlier reports; including them gives +0.88 pp, [0.75, 1.02], but they are not re-derivable from the release and we do not rely on them. These runs are 2,000-step diagnostics on the 0.6B model, and the clean 40k-step rerun confirms the main configuration still fails its utility gate (+0.896, vs 0.9185 on the original cell). The claim that a configuration passes both gates attaches to the four-layer topology, not to the eleven-layer main configuration.
6 When does the structural signal convert to a token advantage? Depth, width, and budget Nine audit cells vary budget, depth, and width (Figure 4; Table 10); each cell is a single run on seed 42, so every shape threshold below is a single-seed reading. Within this run the readings were insensitive to budget between 40k and 100k steps. Depth and width thresholds are read on the paired effect, and every cell’s shuffled-label negative control reads at floor.
paired advantage (pp)
(a) depth ladder
(b) latent width
(c) training budget
1.0
1.0
1.0
0.5
0.5
0.5
0.0
0.0
0.0
−0.5
−0.5
−0.5
split 13 split 17 split 19 (12 delegated) (8 delegated) (6 delegated)
D=64 (headline)
D=96
D=128
40k steps
100k steps
Figure 4: Attack-specific shape across audit cells. (a) depth ladder: the implemented emitter detected the effect at 12–11 delegated layers and not at 8–6. (b) the width sweep: the effect is detected at D=64 and D=96, and not at D=128. (c) The implemented probe’s paired effect did not increase between the two tested exposure budgets. Paired effect over the constant baseline with 95% CI on the gradient arm; one run per cell (seed 42).
13
axis
steps
gradient arm (pp)
negative control (pp)
verdict
depth (12 del.) 2k +0.4796 −0.0702 detected depth (11 del., D=64) 40k +0.9302 −0.0112 detected budget 100k +0.9066 −0.0926 detected width (D=96) 10k +0.8635 −0.0061 detected depth (8 del.) 2k −0.4314 +0.0399 at floor depth (6 del.) 2k −0.2131 +0.0270 at floor width (D=128) 10k −0.4823 −0.0153 at floor width (D=96) 10k −0.0861 +0.0824 at floor width (D=128) 10k −0.1190 +0.0702 at floor (the 512-frame D=96/128 readings are superseded by the 4,096-frame re-runs)
frames 512/512 4,096/4,096 4,096/4,096 4,096/4,096 512/512 512/512 4,096/4,096 512/512 512/512
Table 10: The nine-cell audit. Paired effect of the gradient arm over its constant baseline, with each cell’s shuffled-label negative control beside it; frame-exact counts from the zero-support partition check. All nine committed cells, one run each (seed 42). The partition mechanism (zero-support match) is exact on every cell, whether or not the token effect is detected.
Structural, not statistical. The split is disclosed on every frame because every decoy gradient is exactly zero; the disclosure is deterministic, not a sampling artefact. Exposure budget. The implemented probe’s paired effect did not increase between the two tested exposure budgets. From 40k to 100k gradient exposure (2.5×), the per-row effect changes from +0.9302 to +0.9066 in this run. Attack-specific depth and width dependence. The recovered token advantage is detectable at 12 and 11 delegated layers and undetectable at 8 and 6: the gate is bracketed between 8 and 12 on the matched 2k-step ladder, and the 11-layer point is the 40k-step main configuration, where depth is confounded with budget; the ladder’s point estimates are not a monotone dose–response. The paired effect is detected at latent width D=64 and D=96, and not at D=128. The partition classification itself is exact at every depth and width (Table 10): depth and width affect conversion of the partition into a token advantage for this configuration and attack, not the underlying zero-support signal. Stronger inversion attacks can recover prompts from deep LLM states [8], so this local transition is not evidence that depth generally provides privacy. Table 11 collects every attack family tried against the system, where each applies in the training cycle, and its outcome.
14
attack
step (Fig. 2) outcome
pre-set 9-arm battery on forward frames
6
deep-inversion transformer probe
6
zero-support read (no training; just zeros) 10 gradient arm alone 10 joint arm (frame ∥ gradient)
6+10
XGT cross-Gram invariant
6+10
vector-matching and nearest-signature at- 6 tacks compromise battery (wire capture, de- 6 coy identities, gauge compromise, known plaintext) after the mitigation (per-row clip 0.01 and 10 Gaussian noise)
does it work?
defended floor +0.3.. + 0.9 pp; naked +9.5.. + 10.7 pp
no, against the defence naked split-14 +24.19 pp; defended floor at D=64 and no, against the deD=1024 fence partition exact on every frame, every seed yes floor on the private corpus (+0.50); breaks on 2 of 3 public- only on one dataset corpus seeds (+1.18/ + 1.22) breaks the +1.0 pp gate set in advance on 6 of 6 defended cells yes (+1.44.. + 2.18) floor on every capture (cosine and raw forms) no; the feature carries nothing 0.02–0.03% and 1.5% against a ≈ 5.9% constant baseline, i.e. no below the no-attack control flat at/below control on every fraction no
all arms at floor (+0.48); support 0/1,024; utility cost no longer +0.01 nats
Table 11: The attacker ledger: every attack family tried against the system, where it applies in the cycle of Figure 2, and its outcome. Every committed attack family is listed; the supporting artefacts are indexed in Appendix A. The two leak rows are the finding; the closed row is the fix.
7
External audits and related work
7.1
A selected audit of split-LLM evaluations
The audit selected three recent evaluations to span a bidirectional attack, a bidirectional defence, and recent attack-plus-defence work; it is a purposeful sample, not a field survey (Table 12). None of the three combined all three controls. system
role
injected calibration
shuffled-label negative control
predeclared gate
BiSR [6] DualGuard [20] Prompts to Responses [12]
attack defence attack + defence
none none none
none none no (unmatched baseline)
none none none
Table 12: Selected audit verdicts. The three works report attack or defence performance against reconstruction baselines. “Prompts to Responses” includes an unmatched random-token baseline in its appendix.
BiSR [6] demonstrates backward-channel reconstruction against perturbation defences through comparisons with re-adapted attacks, but does not define an acceptance gate or shuffled-label negative control. DualGuard [20] correctly evaluates forward, backward, and bidirectional attack paths before reporting a worst-case aggregate; its defence comparisons nevertheless do not include an injected detector canary, shuffled-label negative control, or predeclared privacy gate. Thus they do not support the specific calibrated no-leak decision required here; this is a statement about evaluation semantics, not defence efficacy. From Prompts to Responses [12] is forward-only by scope and reports an unmatched random-token baseline in its appendix.
15
7.2
Position within the differential-privacy auditing lineage
Planted-secret exposure tests [4] and attack-based DP audits [15, 22, 29] establish controlled canaries and statistically valid empirical privacy tests. Randomized multi-canary audits already include explicit null hypotheses and decision rules [26]; calibrated membership attacks likewise emphasise reference distributions and declared operating points [3]. We adapt them to channel-level split-protocol evaluation and add a coverage–amplitude dose ladder that estimates each metric’s onset before the gate is applied.
7.3
Split-learning attacks, benchmarks, and adjacent systems
The broader lineage already covers honest-but-curious reconstruction [9], gradient-side label leakage [18], malicious backward-signal control [24], and leakage from intermediate training states [10]. SIMBA [27] and VFLAIR-LLM [11] provide modular evaluation substrates; VFLAIR-LLM reports 5 attacks × 9 defences. Multi-attack split evaluation and bidirectional threat models are therefore well established. The narrower contribution here is the measured positive-control onset, channel-specific shuffled-label negative control, gate set in advance, and artefact-level traceability in one audited system. Table 13 places this report against those systems on the axes that separate them. Within our selected 13-paper audit corpus, five works are in-scope split-LLM evaluations and eight are adjacent systems (trusted execution environment partitioning, inference-time obfuscation, on-device key-value cache protection, prompt sanitisation, and incentive design). MIXGUARD [5] tunes adaptive gradient perturbation against a public proxy and reports defence-vs-attack comparisons, rather than detector calibration against a shuffled-label negative control. Forward-only inference frameworks (e.g., [21, 19, 2]) do not include the training-gradient channel measured here. System
Threat model
This work
compromised cloud; yes: 2k-step cells, forward probe at floor on defended K=3 median 9 declared channels: 6 fresh seeds, 2 cells (conditional: gate fires ≥6% 100%/0% (one 3 measured, 1 con- datasets; 40k leak coverage); joint view breaks the perturbation structible, 5 declared cells gate on 6/6 cells (+1.44..2.18 pp) class) unmeasured until clipping and noise close it (3/3 at floor) honest-but-curious no (inference) point TTRSRs; VMA 13.5–25% on none two models
AloePri [19]
GELO [2] PermLLM [34] ObfNet [32]
VRAM-read + TEE is- no land permutation obfusca- no tion backend holds obfusca- no tors
TEE [35] trust silicon vendor MPC/HE [13, 17, 33, semi-honest parties 16] TOPLOC [23] untrusted provider VFLAIR-LLM [11]
Covers training?
Privacy evidence
p95 cosine/Gram error, no control
scoped out
broken by the matching family none (>99% reconstruction) perceptual (10 volunteers) none
partial no
hardware-rooted attestation cryptographic by assumption
no
n/a (integrity only)
evaluation harness (5 at- substrate tacks × 9 defences)
Integrity
hardware-rooted by construction
Cost 1.05..2.5× WAN
≈0% online; 482 min offline at 671B 20–30% microbenchmark 3 s/token (6B) 0.22–11 ms/sample edge <7–20% 102 –104 ×
hash commits, 258 B / 32 to100%/0% kens shared substrate; its evaluations lack n/a harness the calibrated null
Table 13: Positioning against the closest systems. “Empirical” = attacker-measured privacy; our excess is over a matched no-attack control with a confidence bound and a pre-declared gate, and the thresholds set in advance are calibrated by injected leaks (Section 4). Literature rows report those systems’ own published numbers; their evidence columns use their own authors’ metrics, the uncalibrated, no-matched-null pattern of Section 7, so read the comparison as evaluation method, not as relative leak size.
16
8
Scope and limitations
The findings concern one split-training implementation. The main configuration, calibration, and shape analyses (Section 6) use WikiText-2; a second corpus is a three-seed diagnostic robustness check, not evidence of corpus independence. The external evaluation claim is limited to the targeted comparison of three recent works in Table 12. Five adversarial families are declared unmeasured, not covered: membership/property AUC (needs randomised membership assignment), response-side recovery (no response-side capture), timing metadata (no applicable metric), stateful remote state, and accumulated history. Active perturbation has an applicable calibrated metric but is not executed at the scale of the mitigation runs. The utility side likewise fails on the main configuration (∆ loss 0.9185 on the original cell, 0.896 on the clean rerun, both vs the 0.35 gate). The mitigation runs of Section 5.3 answer that directly: the gate breaks on six of six cells that pass both gates. Those cells are 2,000-step diagnostic runs on the 0.6B model, and convergence-scale evidence remains out of scope. Rare-token recovery is exactly zero on the gradient arm under the implemented emitters (the wire arm shows isolated single-row recoveries, 0.015–0.030%). TAG-, LAMP-, FILM-, or DAGER-style sequence reconstruction has not been adapted to this object, a gradient cut at the split boundary rather than a full parameter gradient [7, 1, 14, 25]. We therefore establish frequent-token discrimination plus exact classification of the constructed partition, not reconstruction of held-out text. Additional negative results illustrate the difficulty of the problem. Rotation alone leaked 18–66% of tokens across three early rotation-only variants; noise strong enough to mask the signal destroyed utility first; longer frames leaked more (+1.25 to +1.36 pp); and a short private phase after public pretraining still leaked (+1.49 pp). A Mutual Information Neural Estimator (MINE) reading near zero nats is not a privacy certificate: a finite-sample lower bound cannot upper-bound true mutual information. The same cell that read −1.3 × 10−6 nats held a detected +0.758 pp probe effect. Two bodies of earlier evidence sit outside the argument above and are recorded in the appendices rather than dropped. Appendix B names each mechanism of the evaluated defence, its implementation site, and its status. Appendix C reports the historical attack batteries against the defended forward cell; they precede the representation-matched positive control of Section 4.2 and are superseded by it as evidence of probe sensitivity. Held-out cross-entropy is measured on blocks held out within one flattened corpus stream rather than on held-out documents (Table 1), so every held-out decision here (the 0.35 utility gate, the claim that the four-layer topology passes it, and the ≈ 0.01 nat mitigation cost) is a block-held-out reading, and claims that depend on document independence are out of scope. On availability: the protocol, metric thresholds, and family map were fixed before the replication and mitigation runs, and those fixed versions are the ones in the artefact release. Committed summaries and code support the displayed calibration and shape results; raw cluster-side prediction tensors are not part of the release. A scorer reimplementation, written internally rather than by an independent group, matched all nine primary seed values to ≤ 10−6 pp; that scorer and its prediction tensors are uncommitted, so the check is recorded but not re-executable. The complete transcript is likewise host-only under the raw-data policy. Appendix A gives the result-by-result artefact index, availability boundaries, recorded verification, and documented limitations.
9
Conclusion
This systems-security case study shows a forward-only evaluation passing while an omitted backward channel revealed which rows were real and which were decoys. The zero-support construction is an implementation and system-design defect; the false pass illustrates the more general evaluation failure mode of omitting an
17
observable channel from the gate. This exact structural metadata disclosure is distinct from the implemented probe’s modest frequent-token paired effect (claim boundaries are recorded in Section 8). On both datasets, all six defended, open-gradient mitigation runs passed the forward privacy and utility gates yet exceeded the privacy gate in the joint view. In the three mitigated runs, per-row gradient clipping plus Gaussian noise removed the zero-support partition signal and suppressed the implemented probes below the gate at a held-out cross-entropy cost of ≈ 0.01 nats. Controlled injection, a predeclared per-metric gate, and shuffled-label negative-control falsification form the calibrated protocol used to evaluate this case. They are not evidence for a general methodology across independently designed systems. In a targeted comparison of three recent evaluations, none combined all three controls. The evidence shows that a passing verdict is meaningful only for the channels, attacks, and leak magnitudes against which its instrument has been tested.
A
Artifact and verification index
All paths in this appendix are relative to the artefact release at https://github.com/setloop-io/Privacy-Failure-in-Split-LLM-Training-The-ReturnedGradient-Nullifies-the-Decoys The index is organised by claim level rather than by directory. A committed summary or manifest supports inspection of the displayed value; “re-derivable” means end-to-end execution from that release alone; availability boundaries are recorded per result below.
A.1
Headline results
A.1.1
Structural metadata disclosure
The exact real/decoy split is recorded for each seed under paper-data/collected/diagnostic/e1_ reproduction_w12/. The primary per-seed bundle summary is named e1_repro_w12_s44_bundles. json for seed 44, with corresponding files for the other seeds. These committed derived records support the frame and row agreement values; the underlying cluster-side prediction tensors are not committed. The complete verified transcript covers 3 seeds, 30,000 optimiser steps, and 180,636 events in the committed manifest. The hardened verifier and the forward, gradient, and joint-view consumers confirm that all 30,000 training frames carry all four payload directions. Payload-level attacks on the accumulated history remain unexecuted, and the raw transcript is host-only. The committed transcript index is under paperdata/collected/diagnostic/w34_complete/. A.1.2
Frequent-token content inference
The exploratory per-seed paired summaries and shuffled-label negative controls are stored with the structural records above. The twelve-seed hierarchical estimate is at paper-data/collected/diagnostic/e1_ hierarchical/ and is generated by bin/summarize_complete_view_matrix.py. The representationmatched positive-control records are under paper-data/collected/diagnostic/deep_probe/; the distinct isolation-audit rerun is under outputs/deep_probe_validation_2026-08-27/. Seed-44 verification trace. The committed per-seed paired record is stored under paper-data/collected/diagnostic/e1_reproduction_w12/. The record is named e1_repro_w12_ s44_arm_grad_real_paired.json. Read best eligible.paired advantage pp: the value +0.6929 18
reproduces Table 8 row-wise. The paired statistic behind that field is computed by bin/paired_advantage. py against the constant baseline over frame-clustered evaluation rows. The adjacent negative-control record, e1_repro_w12_s44_arm_grad_real_shuffled_paired.json, must read at floor. This is the recorded verification trace for the reported value. A.1.3
Calibration and attack-specific shape
The declared-channel table uses paper-data/family_metric_map.json together with the mitigation-run results, and is regenerated with bin/build_channel_table.py. The dose–response source is paperdata/collected/diagnostic/w24_metric_sweep/w24_dose_response.json, the associated summaries are under paper-data/collected/diagnostic/w24_metric_sweep/, and the figure command is papers/paper-1/figs/build_figures.py. The threshold map is paper-data/family_metric_ map.json. The nine-cell shape records comprise the paper-data/collected/diagnostic/gradaudit/ cell JSONs. Their full summary is paper-data/collected/diagnostic/gradaudit/w56_gradaudit_summary.json. They use the same papers/paper-1/figs/build_figures.py command. These calibration and shape rows are re-derivable from committed summaries and code.
A.2
Confirmation
A.2.1
The replication runs
Seeds 48–50 are retained under paper-data/collected/diagnostic/e1_confirmation/. Their committed paired summaries support the displayed values; raw prediction tensors remain cluster-side. A.2.2
The mitigation runs
The mitigation-run cells, paired statistics, and SHA-256 manifest of the 7.4 GB bundle store are committed under paper-data/collected/diagnostic/phasec_2026-08-27/. Per-cell records are named paired_summary.json. The packaged driver and scorer redisplay the committed derived summaries; the raw bundle store remains on the cluster. A.2.3
Internal scorer availability
The independent scorer check is recorded in the internal red-team review of the scorer reimplementation (docs/audits/W54_RED_TEAM_2026-08-26.md). It re-derived all nine primary seed values to ≤ 10−6 pp. The report is committed, but the independent scorer and its recorded prediction tensors are not, so this verification is recorded but not repository-re-executable. Table 14 consolidates the availability boundary for these result categories.
19
result category
committed input
committed code
re-derivable
exploratory gradient result nine audit cells
packaged scorer
displayed values from committed summaries yes
mitigation runs
per-seed paired JSONs and negative controls; bundles are manifests to cluster-side prediction tensors paper-data/collected/diagnostic/gradaudit/ cell JSONs and summary paperdata/collected/diagnostic/w24_metric_sweep/ JSONs phasec 2026-08-27/ manifests and paired summaries
internal scorer check
internal-check report
complete transcript
hashes and manifest
dose–response calibration
papers/paper1/figs/build_figures.py papers/paper1/figs/build_figures.py packaged driver and scorer uncommitted independent scorer verifier
yes
displayed values from committed summaries recorded verification only no; host-only by policy
Table 14: Per-result availability. The audit-cell and calibration rows regenerate from committed summaries; the exploratory and mitigation rows are redisplayed from committed derived artefacts whose raw prediction tensors remain on the cluster. Transcript hashes and the manifest are committed, but the payload remains on its host.
A.3
Corrections and limitations
A.3.1
Sequential-block split versus the document protocol
The paper-data/evaluation_protocol.json file, fixed in advance, declares a document-level held-out split, but the committed runner implements the sequential fixed-width block split reported in Table 1. Held-out measurements are therefore block-held-out within one flattened corpus stream, not document-held-out; claims that depend on document independence remain out of scope. This records the discrepancy; the protocol file itself was left unedited. A.3.2
Corrections of record
The traceability verification is recorded at the traceability audit, which fixes the committed-versus-cluster-side evidence boundary (docs/audits/W78_TRACEABILITY_2026-08-26.md). It supersedes earlier wording about what is committed versus cluster-side. The adversarial review round, an external adversarial re-read of the manuscript’s evidence claims, at docs/audits/W78_CODEX_REVIEW_2026-08-27.md applies it, and the same red-team report records the independent re-derivation. The traceability audit has precedence for evidence boundaries. A forward-membership reading of +0.068 (AUC above the 0.5 floor, stable across all six packaged seeds) was falsified by the shuffled-label negative control: it survives label permutation and randomised membership assignment, decomposes into in-sample memorisation (+0.095) with a corpus-region term of ≈ 0, and exists only in the coordinate probe that cannot generalise. The purported members were the probe’s own training rows. We withdrew the channel; and membership auc remains a diagnostic. The underlying capture remains on the cluster. Claims withdrawn during the study are recorded in paper-data/claim_evidence_ledger.json with their refutation records; the appendices that follow retain withdrawn entries only to identify their status and provenance.
B
The defended stack, mechanism by mechanism
Table 15 names each mechanism of the evaluated defence concretely, with its implementation site and its status in the current evidence. Three rows carry a non-core status, each for a recorded reason: the fragmentation cell’s remote modules never received a gradient (a dispatch defect; the cell is invalidated 20
as a trained-fragmentation experiment), the public-pretraining and capacity framing are withdrawn per the refutation record, and the deep defended cloud’s capacity conclusion is withdrawn with it. Purpose
Technique
Detail
Status
adversarially trained MLP en- H→D=64 MLP with GELU, trained by minimax against embedded core coder/decoder (information token/property/reconstruction probes plus a MINE MI penalty (the MI bottleneck) reading is not a certificate) per-token-row L2 clip to C=1.0 + Gaussian mechanism with zCDP accounting (the ε is quoted as vacuous) core Noise additive Gaussian σ=0.35·C (0.40 on hybrid MoE), both directions Key derivation SHA-256 counter-mode KDF 128-bit CSPRNG master per draw, domain-separated core sha256(DOM∥S∥t), exact (u64≫11)·2−53 uniforms, float64 Box-Muller Rotation Haar-random orthogonal Gaussian matrix → QR with positive-diagonal sign fix, at D width; core fresh per request Permutation Fisher–Yates over the KDF stream without replacement; fresh per request core Scale gauge per-row exp(N (0, σ)) · (±1), independent of the other gauges (separate KDF epoch); OFF for core σ=0.75, clamped [0.2, 5] radial/invariant-MLP clouds recycled real latent rows from ear- CSPRNG-sampled without replacement, honest labels tracked; decoys core Decoy rows lier train blocks are real data, not noise Transport websockets over TLS 1.3 pinned self-signed CA (trusts no system store), hostname verification core on, plaintext refused Session isolation fresh cloud model + AdamW per seeded per session under fork rng; no cross-session state core connection Gradient harden- per-token L2 clip 1.0, non-finite ⇒ on remote-returned gradients core ing abort Integrity identical-seed replicas + median relative deviation from group mean > 0.02 ⇒ flagged; median neu- avail. compare tralises one adversary; demonstrated on one large perturbation class only gauge-equivariant Gram message softmax over squared unit-Gram × Gram @ unit rows (monomial), core Cloud compute passing expert routing, or invariant-feature MLP gates Wire quantisation fixed-grid int8 over the gauge clamp straight-through on the defender gradient path; absmax variant rejected avail. range (strips the scale gauge) Replay resistance per-request nonce acceptance win- a replayed frame is refused core dow Key ratchet per-epoch key derivation on the limits cross-request capture pooling; the current transform is necessarily core KDF chain visible to the executor; it bounds accumulation, it is not irreversibility Redundant chan- K=2 independent cloud channels, cross-check + independence core nels fresh gauges per channel Two-segment del- cloud segment A / private island / chained surrogate error compounds across segments, the measured opt. egation cloud segment B reason it stays optional Public pretraining UCN trains on public data at full the short-private-phase cell leaked (+1.49 pp); privacy cost zero on the opt. fidelity; private phase re-hardens public stream Fragmentation frame split into two D=32 half- structural partition across servers; the measured cell’s remote modules invalid. views never trained, invalidating it as a trained-fragmentation experiment Gram flattening regularizer penalising off-diagonal keeps the invariant features uninformative avail. Gram concentration Deep defended deep invariant-feature MLP: norms 118.9M params; its capacity conclusion is withdrawn withdr. cloud (isolation and Gram-statistic feature channel cells) with deep wide gates Complementary K=3 disjoint 64-wide views, one privacy rests on k-of-K collusion; union attacker evaluated conserva- opt. views per cloud, per-view rotation, one tively shared row permutation Secret mixup released rows = λz + (1−λ)zdecoy , TLN corrects the return via a parallel decoy forward; approximate under opt. λ=0.75 the nonlinear cloud, cost measured Adaptive con- fresh probe every 100 steps; trip- one gross breach repaired within one window; its monitor was later withdr. troller wire actions + utility-gate width lad- shown blind at relevant operating points der Compression
Table 15: Techniques at a glance: what each mechanism is, concretely, where it lives, and its status (core / available / optional / withdrawn / invalidated) in the current evidence.
None of these are novel primitives in isolation. QR rotations, Fisher–Yates, Gaussian mechanisms, TLS are all standard. The engineering content is their composition order and their measured interaction: the 21
bottleneck alone leaks, the gauges alone leak, noise alone destroys utility. The composite at these settings is what sits at the floor. That is why the table names not just the technique but the configuration it was validated at.
C
The defended cell under its historical attack batteries
The three batteries below attack the defended forward cell (the conditional claim), not the backward channel this report is about. Each uses one probe family on the forward view only, and the battery’s convention reads trends below a weak constant baseline. They are reported here with the representation-matched positive control (Section 4.2, Table 6) that the original batteries lacked. These batteries were run on two models: the Qwen3-0.6B of Table 1, and a 35B-A3B mixture-of-experts model used only here, as a capacity check on the defended forward cell. No result in the body of this report depends on the larger model. Table 16 reports the compromise-fraction battery, Table 17 the external matching-attack families, and Table 18 the Byzantine median verification. Arm (fraction grid)
0.6B trend
35B-A3B trend
Reading
Wire capture 10/25/50/100% Decoy identities revealed 0/25/50/90%
flat, −4.7 to −4.9 −4.9 to −7.4
flat, −4.6 to −5.1 −4.9 to −5.6
Gauge secrets leaked, 10–20%
−7.8 ungauged / gauged flat, −3.8 to −4.4
capturing the whole wire does not help knowing which rows are decoys makes the probe worse rows seen before gauging still read at guessing level knowing some tokens does not reveal the others
Known plaintext 1/5/20%
−4.7
−5.7 / −4.6 flat, ≈ −4.0
Table 16: Compromise-fraction battery on the defended cell. Each row gives the attacker a growing share of some secret and asks whether the probe improves. Values are the probe’s accuracy in percentage points above or below its constant-guess baseline, so a negative number means the probe did worse than guessing. None of these arms approaches the +1.0 pp gate, and none improves as the attacker is given more.
Model 0.6B (seed 42) 35B-A3B (seed 42)
Gate statistic (pp)
Vector-matching top-1, against baseline
Nearest-signature top-1, against baseline
+0.541 +0.285
0.03% vs 5.92% 0.02% vs 5.85%
1.56% vs 5.92% 1.54% vs 5.85%
Table 17: External matching-attack families vs. the defended cell (regenerated bundles). Two adapted attack families on these captures. The positive-control sensitivity they presuppose is Table 6.
Cell 3 honest CPU nodes 2 honest + 1 malicious (lying)
Verified frames
Flagged
Max rel. dev.
Loss ∆
Ratio
2,256 2,256
0 (0% FP) 2,256 (100%)
4.98×10−8 0.5947
+0.071 +0.088
9.30× 9.13×
Table 18: K=3 Byzantine median verification. One large perturbation class on identical CPU replicas; no general active-security claim.
References [1] Mislav Balunović, Dimitar I. Dimitrov, Nikola Jovanović, and Martin Vechev. LAMP: Extracting text from gradients with language model priors. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022.
22
[2] Anatoly Belikov and Ilya Fedotov. Good-enough LLM obfuscation, 2026. [3] Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramèr. Membership inference attacks from first principles. In 2022 IEEE Symposium on Security and Privacy, pages 1897–1914, 2022. [4] Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th USENIX Security Symposium (USENIX Security 19), pages 267–284. USENIX Association, August 2019. [5] Chen Chen, Xiang Gao, Xianshun Wang, Chengran Li, Shengyu Xia, Xueluan Gong, Linru Zhang, Qian Wang, and Kwok-Yan Lam. The art of mixology: Mixup-based obfuscation for privacy-preserving split learning in large language models, 2026. [6] Guanzhong Chen, Zhenghan Qin, Mingxin Yang, Yajie Zhou, Tao Fan, Tianyu Du, and Zenglin Xu. Unveiling the vulnerability of private fine-tuning in split-based frameworks for large language models: A bidirectionally enhanced attack. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security, pages 2904–2918. Association for Computing Machinery, 2024. arXiv:2409.00960. [7] Jieren Deng, Yijue Wang, Ji Li, Chenghong Wang, Chao Shang, Hang Liu, Sanguthevar Rajasekaran, and Caiwen Ding. TAG: Gradient attack on transformer-based language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3600–3610. Association for Computational Linguistics, November 2021. [8] Tian Dong, Yan Meng, Shaofeng Li, Guoxing Chen, Zhen Liu, and Haojin Zhu. Depth gives a false sense of privacy: LLM internal states inversion. In 34th USENIX Security Symposium (USENIX Security 25), pages 1629–1648. USENIX Association, August 2025. [9] Ege Erdogan, Alptekin Küpçü, and A. Ercüment Çiçek. UnSplit: Data-oblivious model inversion, model stealing, and label inference attacks against split learning. In Proceedings of the 2022 Workshop on Privacy in the Electronic Society (WPES), pages 115–124, 2022. [10] Xinben Gao and Lan Zhang. PCAT: Functionality and data stealing from split learning by pseudo-client attack. In 32nd USENIX Security Symposium (USENIX Security 23), pages 5271–5288. USENIX Association, August 2023. [11] Zixuan Gu, Qiufeng Fan, Long Sun, Yang Liu, and Xiaojun Ye. VFLAIR-LLM: A comprehensive framework and benchmark for split learning of LLMs. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, pages 5470–5481. Association for Computing Machinery, 2025. arXiv:2508.03097. [12] Zixuan Gu, Xiaojun Ye, and Yang Liu. From prompts to responses: Dual-sided data leakage and defense in split large language models, 2026. [13] Kanav Gupta, Neha Jawalkar, Ananta Mukherjee, Nishanth Chandran, Divya Gupta, Ashish Panwar, and Rahul Sharma. SIGMA: Secure GPT inference with function secret sharing. PoPETs, 2024(4), 2024. [14] Samyak Gupta, Yangsibo Huang, Zexuan Zhong, Tianyu Gao, Kai Li, and Danqi Chen. Recovering private text in federated learning of language models. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022. 23
[15] Matthew Jagielski, Jonathan Ullman, and Alina Oprea. Auditing differentially private machine learning: How private is private SGD? In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 2020. [16] Wen jie Lu et al. PUMA: Secure inference of LLaMA-7B in five minutes, 2023. [17] Wen jie Lu et al. BumbleBee: Secure two-party inference framework for large transformers. In NDSS, 2025. [18] Oscar Li, Jiankai Sun, Xin Yang, Weihao Gao, Hongyi Zhang, Junyuan Xie, Virginia Smith, and Chong Wang. Label leakage and protection in two-party split learning. In International Conference on Learning Representations (ICLR), 2022. [19] Yu Lin, Qizhi Zhang, Wenqiang Ruan, Daode Zhang, Jue Hong, Ye Wu, Hanning Xia, Yunlong Mao, and Sheng Zhong. Towards privacy-preserving LLM inference via covariant obfuscation, 2026. v2; v1 title used “Collaborative”. [20] Zihan Liu, Yizhen Wang, Rui Wang, and Sai Wu. DualGuard: A parameter space transformation approach for bidirectional defense in split-based LLM fine-tuning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17065–17080, Vienna, Austria, July 2025. Association for Computational Linguistics. [21] Xinjian Luo, Ting Yu, and Xiaokui Xiao. Prompt inference attack on distributed large language model inference frameworks. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, pages 1739–1753, 2025. arXiv:2503.09291. [22] Milad Nasr, Jamie Hayes, Thomas Steinke, Borja Balle, Florian Tramèr, Matthew Jagielski, Nicholas Carlini, and Andreas Terzis. Tight auditing of differentially private machine learning. In 32nd USENIX Security Symposium (USENIX Security 23), pages 1631–1648, Anaheim, CA, August 2023. USENIX Association. [23] Jack Min Ong, Matthew Di Ferrante, Aaron Pazdera, Ryan Garner, Sami Jaghouar, Manveer Basra, Max Ryabinin, and Johannes Hagemann. TOPLOC: A locality sensitive hashing scheme for trustless verifiable inference, 2025. [24] Dario Pasquini, Giuseppe Ateniese, and Massimo Bernaschi. Unleashing the tiger: Inference attacks on split learning. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, pages 2113–2129, 2021. [25] Ivo Petrov, Dimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller, and Martin Vechev. DAGER: Exact gradient inversion for large language models. In Advances in Neural Information Processing Systems (NeurIPS), volume 37, 2024. [26] Krishna Pillutla, Galen Andrew, Peter Kairouz, H. Brendan McMahan, Alina Oprea, and Sewoong Oh. Unleashing the power of randomization in auditing differentially private ML. In Federated Learning and Analytics in Practice Workshop at ICML, 2023. [27] Abhishek Singh, Vivek Sharma, Rohan Sukumaran, John Mose, Jeffrey Chiu, Justin Yu, and Ramesh Raskar. SIMBA: Split inference—mechanisms, benchmarks and attacks. In Computer Vision – ECCV 2024, volume 15134 of Lecture Notes in Computer Science, pages 214–232. Springer Nature Switzerland, 2024.
24
[28] Liwei Song and Prateek Mittal. Systematic evaluation of privacy risks of machine learning models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2615–2632. USENIX Association, August 2021. [29] Thomas Steinke, Milad Nasr, and Matthew Jagielski. Privacy auditing with one (1) training run. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023. Full version: arXiv:2305.08846. [30] Xiaoxiao Sun, Nidham Gazagnadou, Vivek Sharma, Lingjuan Lyu, Hongdong Li, and Liang Zheng. Privacy assessment on reconstructed images: Are existing evaluation metrics faithful to human perception? In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023. [31] Praneeth Vepakomma, Otkrist Gupta, Tristan Swedish, and Ramesh Raskar. Split learning for health: Distributed deep learning without sharing raw patient data, 2018. [32] Rui Xu et al. Lightweight and unobtrusive data obfuscation at IoT edge for remote inference, 2019. [33] Jiawen Zhang et al. NEXUS: Secure transformer inference made non-interactive. In NDSS, 2025. IACR ePrint 2024/136. [34] Fei Zheng, Chaochao Chen, Zhongxuan Han, and Xiaolin Zheng. PermLLM: Private inference of large language models within 3 seconds under WAN, 2024. [35] Jianwei Zhu, Hang Yin, Peng Deng, Aline Almeida, and Shunfan Zhou. Confidential computing on NVIDIA hopper GPUs: A performance benchmark study, 2024. [36] Ligeng Zhu, Zhijian Liu, and Song Han. Deep leakage from gradients. In Advances in Neural Information Processing Systems (NeurIPS), volume 32, 2019.
25