Conceptio › Archive › arXiv CS
arXiv CSopen access

Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs

arXiv:2609.20532v1 [cs.CR] 17 Sep 2026

Li Ge1,* , Wenjie Qu2,* , Weitao Feng1 , Yi Zeng1 , Jiaheng Zhang2 , Xiaofeng Wang1 , Wei Dong1 1 Nanyang Technological University 2 National University of Singapore * Equal contribution.

Abstract

ference, and attribute exposure, have become a serious concern [9, 15]. The dominant mitigation is Differential Privacy (DP) [13], which ensures that any single record has only a limited effect on an algorithm’s output. In ML training, DP is commonly implemented through DP-SGD [1], which adds calibrated noise to per-example gradients before aggregation, so the final model satisfies DP. DP and DP-SGD have thus been increasingly recommended by privacy laws, regulations, and best-practice guidelines for model protection [3, 29, 35]. Verifiable DP training. However, assuring that DP-SGD is faithfully enforced during model training remains challenging. The key issue is verifying the integrity of the training process: whether the declared DP mechanism has been executed correctly throughout training without deviation. This is critical because the noise injected at each iteration inevitably degrades model utility, creating incentives for model owners to deviate from the protocol while still claiming privacy protection. For example, a company subject to regulatory requirements or public commitments may reduce or even omit DP noise to obtain a higher-performing model. Importantly, this concern is not merely hypothetical. Several real-world deployments of DP have attracted scrutiny regarding whether their implementations provide the same guarantees as claimed [39]. A recent work reports that one building block in Apple’s DP framework in real-time is configured with DP disabled, thus uploading data without any DP protection [10]. Although extensive research has focused on designing DP mechanisms [14, 24] and analyzing or auditing their privacy guarantees [19], little attention has been paid to verifying that DP-SGD is actually executed as claimed, despite the importance of such assurance for privacy protection and regulatory compliance. Addressing this concern requires verifiable DP [8, 26], which aims to provide publicly checkable evidence that a claimed DP mechanism has been executed as specified. Given such evidence, together with the declared protocol and the resulting model, an independent verifier should be able to determine that the model was produced under the attested verification protocol and its declared probabilistic security bounds. Importantly, this verification itself must preserve pri-

Wide adoption of machine learning has created growing policy and regulatory demand for protecting sensitive training data, with differential privacy (DP) emerging as a key mechanism. Yet a less-studied problem is how to certify the faithful execution of DP during training: an external verifier should be able to check that a released model was trained with proper DP protection, without accessing the private training data. Existing cryptographic approaches, such as zero-knowledge proofs, provide strong guarantees but often incur prohibitive overhead, in some cases by orders of magnitude. Trusted Execution Environments (TEEs) offer a more efficient alternative, but the multi-GPU TEE support needed for training and fine-tuning large language models remains limited to recent platforms and is absent or inefficient on legacy GPUs. To address this, we propose a practical framework for verifiable DP training using CPU-side TEEs together with untrusted GPUs. Our design addresses a fundamental efficiencysecurity tension: training entirely inside a CPU TEE is too slow, while unrestricted GPU offloading can allow malicious deviations from DP. We therefore offload expensive gradient computation to GPUs, while using the CPU TEE to efficiently verify the correct enforcement of DP on gradients through probabilistic checking. Our framework detects frequent full deviations from DP with high probability; for the utility-oriented forged-gradient attacks evaluated in this work, sparse deviations provide limited utility benefit and show no measurable additional membership leakage. Experiments further show that our approach nearly achieves a “free lunch”: it incurs only modest overhead compared with standard GPUbased DP training, while effectively constraining malicious deviations from the claimed DP execution.

1

Introduction

As sensitive data such as electronic health records and tax records are increasingly used in machine learning (ML) training, privacy risks, including reidentification, membership in1

vacy and should not require access to sensitive training data. In the example above, after a company trains a model with a claimed DP guarantee, a regulator or external auditor may require evidence showing that the declared DP mechanism was actually enforced during training. Verifiable privacy mechanisms have been studied for tasks such as k-means clustering [26], DP counting, and histogram release [8,26]. Extending such guarantees to deep model training, however, remains challenging. Existing state-of-the-art approaches [34] rely on zero-knowledge proofs (ZKPs) to certify the correctness of the training procedure. While providing strong guarantees, ZKPs incur substantial computational and communication costs that grow with model size and the number of training iterations. Even for relatively small models with only 0.01B parameters, the overhead can exceed 1000×, making these approaches impractical for realistic deep learning workloads. TEE-based alternative. Beyond heavyweight ZKP-based solutions, a natural alternative is to use hardware-based trusted execution environments (TEEs). A TEE provides an isolated execution environment whose code identity and execution state can be remotely attested. This suggests a simple design: run the entire training procedure inside a TEE and use the hardware root of trust to certify, through an open-source monitor, that the DP protocol was correctly executed. An external verifier can then check both that the declared DP protocol was faithfully enforced and that the released model is bound to the attested execution. However, this “train-inside-a-TEE” design faces a practical obstacle: training reasonably sized models requires GPUs. Full confidential-computing support across both CPU and GPU TEEs is available only on the most recent platforms, such as NVIDIA Blackwell B200 [28]. Older GPUs, including H100 and H200 based on the Hopper architecture, rely on secure I/O through encrypted bounce buffers, which introduces substantial training overhead [5, 49]. Earlier-generation GPUs, such as V100, A100, and RTX 4090, provide no TEE support at all. Since GPU infrastructure is costly and replaced slowly, it is unrealistic to assume that GPU-TEEs will be widely available in the near future [44]. Our approach. We therefore propose a new approach: combining widely deployed CPU-TEEs with untrusted GPUs. CPU-TEEs, such as AMD SEV [33] and Intel TDX [4], are already available on commodity servers, but are too slow to run full training by themselves (Section 8). Our key idea is instead to use the CPU-TEE to monitor DP-SGD executed on untrusted GPUs and generate verifiable evidence of protocol compliance. This design preserves the efficiency of GPU-based training while providing verifiable DP guarantees with only modest overhead over conventional, non-verifiable DP-SGD. More specifically, we design and implement a CPU-TEEbased framework for verifiable DP training that incurs only modest overhead compared with standard GPU-based DP-

SGD. In the absence of GPU TEE support and efficient secure I/O, our framework achieves high performance by offloading expensive gradient computation to untrusted local GPUs while keeping trusted randomness, optimizer and model-state maintenance, and verification inside the CPU TEE; the GPU performs per-example gradient computation, clipping, and aggregation, subject to probabilistic trusted recomputation. This architecture, however, introduces a challenge: an untrusted trainer may tamper with GPU-computed gradients before they enter the TEE, causing the resulting model to deviate from the claimed DP protocol (Section 4). To address this challenge, we classify malicious deviations according to their frequency, distinguishing between dense and sparse deviations. To understand how much utility a trainer can obtain from infrequent manipulation, we study several utility-oriented forged-gradient attacks that steer the DP trajectory toward non-private checkpoints. Across the attack family and workloads evaluated in this paper, small attack budgets provide limited utility gain and no measurable additional membership leakage, leaving a rational trainer with limited incentive to employ them; we use these experiments to select an empirical sparse–dense operating threshold rather than to claim that every possible sparse attack is ineffective. Motivated by this observation, we design an efficient CPUTEE-based protocol that probabilistically detects repeated full deviations while explicitly accounting for the cumulative manipulation admitted through its numerical-tolerance region. The GPU performs per-example gradient computation, clipping, and aggregation, whereas the CPU-TEE holds the trusted randomness and training state, generates the DP noise, and probabilistically verifies the GPU-side computation. As a result, frequent deviations are detected with high probability, while the impact of infrequent deviations remains bounded by DP-SGD clipping, enabling efficient verifiable DP training under realistic hardware assumptions. Contributions. Our contributions are listed below: • System Design: We propose a CPU-TEE-based solution for verifiable DP training that is compatible with legacy GPU infrastructure. We identify an inherent efficiency–security tension in this setting. To address this tension, we design a framework that achieves a practical trade-off between security and efficiency. Our framework offloads gradient computation, per-example clipping, and aggregation to untrusted GPUs, while keeping trusted randomness, model and optimizer state, DP-noise generation, and probabilistic verification inside the CPU-side TEE. The protocol probabilistically detects repeated full deviations while explicitly accounting for the cumulative manipulation admitted through its numericaltolerance region, and avoids the cost of fully verifying every training step. • Adversarial Analysis: We study three utility-oriented checkpoint-steering attacks—Current, Nearest, and Final— under adversarially selected attack schedules. Across the eval2

uated workloads, attack budgets below M = 50 malicious iterations (0.33%–0.89% of the training steps on the LLM workloads, 1% on CIFAR-10) provide limited or no utility improvement, and our membership-inference evaluation shows no measurable additional leakage over the honest DP baselines. These findings motivate M = 50 as an empirical operating threshold for the attack family considered here; they do not claim to exhaust all possible utility- or privacy-oriented attacks. • Numerical Discrepancy as an Attack Surface: We identify that the numerical discrepancy between honest GPU and TEE executions is itself an attack surface: any tolerance the verifier must grant for honest disagreement can be exploited by a dishonest GPU to hide bounded manipulations that no single check would flag. Independently of the attack strategy, our two-stage verification protocol classifies every committed submission into a body, ambiguity, or hard-rejection region, meters tolerated discrepancies observed on checked steps with bounded ledgers, and yields an analytic high-probability security budget for the cumulative normalized deviation that can pass through the first two regions, alongside a Bernoulli detection guarantee for the third. • Experimental Results: We conduct a comprehensive empirical evaluation across multiple models, datasets, and training configurations. Our experiments show that verifiable DP training is nearly a “free lunch” on LLM oriented tasks: our framework incurs only 4.4%–14.9% overhead compared with standard GPU-based DP training across a range of tasks and datasets. In contrast, executing the full training procedure inside a CPU-side TEE incurs 8.0×–9.2× overhead.

as the model produced by the randomized training algorithm on dataset D .

2.2

DP-SGD

Differentially Private Stochastic Gradient Descent (DPSGD) [1] enforces differential privacy during training by bounding the influence of each individual example through per-example gradient clipping and injecting Gaussian noise into the aggregated gradient. Clipping prevents any single record from dominating the update, while the added noise masks the contribution of individual examples and provides a formal privacy guarantee. Let D = {xi }Ni=1 be the training dataset, θ the model parameters, and ℓ(θ; x) the loss for an example x. At training iteration t, given clipping norm C, noise multiplier σ, and learning rate ηt , the per-example gradient on xi is gt (xi ) = ∇θ ℓ(θt ; xi ). Each gradient is clipped to have ℓ2 norm at most C:   C ḡt (xi ) = gt (xi ) · min 1, . ∥gt (xi )∥2 DP-SGD then aggregates the clipped gradients in the training batch Bt and adds Gaussian noise: !  1 e gt = ḡt (xi ) + zt , zt ∼ N 0, σ2C2 I . ∑ |Bt | xi ∈Bt Finally, the model parameters are updated as

2 2.1

Preliminaries

θt+1 = θt − ηt e gt .

Differential Privacy

2.3

Differential Privacy (DP) [13] is a rigorous framework for quantifying the privacy protection provided by a randomized algorithm. Intuitively, DP requires that the output of an algorithm be only minimally affected by the presence or absence of any individual data record, thereby limiting the information that can be inferred about a particular participant. In this work, we adopt the add/remove-one notion of neighboring datasets. Two datasets D and D ′ are said to be neighboring, denoted by D ∼ D ′ , if they differ by the addition or removal of a single record. Given privacy parameters ε, δ ≥ 0, a randomized algorithm M is said to satisfy (ε, δ)-DP if for any pair of neighboring datasets D ∼ D ′ and any measurable set of outputs S ,

Trusted Execution Environment

A Trusted Execution Environment (TEE) is a hardwarebacked isolated execution context that provides confidentiality and integrity for code and data during execution, even if the host operating system or hypervisor is compromised. Typical TEEs (e.g., Intel SGX, AMD SEV-style enclaves) offer attestation mechanisms: a remote party can verify that a specific program (identified by a measurement such as a hash) is running inside a genuine TEE before trusting its outputs. In our setting, attestation enables an auditor to trust that the DP-enforcing logic (noise generation, model updates, and the recomputation that checks GPU-side clipping and aggregation) is executed as specified. However, due to legacy deployment constraints discussed in the introduction, most widely available TEEs today are CPU-side TEEs, which impose practical limitations. Enclave transitions can be expensive, and full-scale training entirely within a TEE is often impractical—especially for large models that rely on high-throughput GPU computation. Therefore, we

Pr[M (D ) ∈ S ] ≤ eε Pr[M (D ′ ) ∈ S ] + δ. Therefore, DP requires the output distributions of M on neighboring datasets to be close, thereby limiting privacy leakage. In the context of private model training, M (D ) can be viewed 3

treat the TEE as a trusted core that holds the trusted randomness, the optimizer and model state, and the verification logic, and records verifiable evidence, while outsourcing heavy gradient computation to an untrusted but fast GPU worker whose results are checked by probabilistic trusted recomputation. Unless otherwise specified, all TEEs are referred to as CPUTEEs throughout this document.

3

The regulator does not need to observe every low-level operation. Instead, it requires that privacy-critical parts of the protocol be either protected by trusted hardware or auditable afterward. In this sense, the TEE serves as a bridge between the company and the regulator: it protects key DP operations and helps generate verifiable evidence that the claimed protocol was enforced. Users therefore rely on the regulator’s compliance framework rather than directly observing the training process. Our goal is to achieve verifiable DP training, where the training procedure generates a post-training certificate that can be released to users or used for future attestation. Such a design constrains the company from deviating from the prescribed protocol. As a result, the company is expected to faithfully follow the DP-SGD procedure, users obtain meaningful privacy protection, and regulators or users receive verifiable evidence rather than relying on unaudited privacy claims.

Problem Definition and Security Model

In this work, we consider a three-party setting that captures the practical deployment of verifiable DP training. User (data provider). The user contributes data records, such as samples and labels, to form the training dataset. Users care about privacy: they require a formal DP guarantee that bounds each participant’s privacy leakage and limits what can be inferred about their individual records from the released model. Company (trainer). The trainer (e.g., a company operating the training pipeline and GPU workers) controls the model architecture, training pipeline, and computational resources. The trainer aims to obtain a high-utility model efficiently. However, because differential privacy typically reduces model utility, a profit-driven or malicious trainer may have incentives to deviate from the prescribed DP-SGD protocol while still claiming compliance to users or regulators. In particular, the trainer may attempt to improve model utility by modifying key components of DP-SGD. Typical deviations include: (i) skipping or weakening gradient clipping, thereby increasing sensitivity; (ii) reducing the noise scale σ or replacing Gaussian noise with a weaker distribution; (iii) reusing noise across steps or employing predictable randomness; (iv) altering the sampling procedure, for example by using larger effective batch sizes than declared; and (v) selectively applying DP mechanisms only to a subset of training steps. Government (regulator/auditor). The government (or regulator) is an external party that enforces compliance with privacy requirements. The regulator does not necessarily participate in training, but it demands verifiability: the company should provide convincing evidence that DP-SGD was executed correctly with the claimed hyperparameters (e.g., sampling rate, clipping norm, and noise multiplier). Interaction among the three parties. The interaction proceeds as follows. Users provide data to the company, expecting that the model will be trained under the advertised DP guarantee. The company then trains the model using its own infrastructure, which may include both untrusted highperformance components, such as GPUs, and trusted components, such as a CPU-based TEE. During training, the company is expected to follow the declared DP-SGD protocol and produce evidence that the regulator can later use to assess compliance.

4

Challenges of Verifiable DP Training

In this section, we identify the challenges of realizing verifiable DP training under realistic hardware assumptions. We begin by considering two natural approaches and show that each is limited by either efficiency or security concerns. Together, these limitations reveal a fundamental efficiency-security tradeoff that motivates our design. TEE-only training. A straightforward way to achieve verifiable DP training is to execute the entire training process inside a TEE. Under this design, model parameters, gradient computation, clipping, noise generation, and parameter updates are all performed within the trusted boundary. Through remote attestation, a verifier can therefore obtain strong assurance that the declared DP-SGD protocol has been faithfully executed. Efficiency limitation. Although conceptually simple, this approach is impractical for modern learning workloads. Largescale training, such as training a large language model, relies heavily on GPU acceleration, whereas CPU-based TEEs provide substantially lower computational throughput. As a result, repeatedly performing forward and backward propagation inside a CPU-TEE incurs prohibitive overhead. As shown in Section 8, this design slows training by approximately 8.0×–9.2× on the evaluated LLM workloads and by about 20× on CIFAR-10 compared with standard GPU-based DP training. Consequently, fully executing DP-SGD inside a CPU-TEE is too expensive for realistic deployments. GPU-TEE split training. To improve efficiency, a natural alternative is to offload computationally intensive operations from the trusted environment to untrusted local hardware. In this design, the GPU performs the expensive forward and backward passes, while the TEE maintains the trusted model state and executes the privacy-preserving update. Concretely, at each iteration, the GPU computes the gradients for the 4

current mini-batch and sends them to the TEE. The TEE then performs the DP-SGD update, including clipping and noise addition, updates the trusted model parameters, and then returns the updated state information to the local GPU for the next iteration. Security limitation. While substantially more efficient, this design introduces a critical security gap. The TEE only observes the gradients received from the GPU and cannot directly verify whether they were produced by honest forward and backward propagation on the intended model and training batch. Consequently, a malicious trainer may manipulate GPU-side computation while continuing to interact with the TEE in a seemingly legitimate manner. More concretely, the GPU may return gradients that do not correspond to the claimed training process. By carefully crafting such gradients, the trainer can steer optimization toward a higher-utility model while violating the integrity of the declared DP-SGD execution. For example, the trainer may maintain two models in parallel: a DP-compliant model that remains consistent with the TEE’s view and a non-private model trained outside the trusted environment. During training, manipulated gradients can gradually reduce the discrepancy between these two models, allowing the final model to benefit from nonprivate training while still appearing to follow the declared DP protocol.

ited utility benefit and no measurable additional membership leakage.

5.1

As discussed in the previous section, verifiable DP training exhibits an inherent tension between security and efficiency. The two straightforward solutions represent opposite ends of this tradeoff. Executing the entire training procedure inside a TEE can, in principle, eliminate adversarial behavior, but incurs substantial computational overhead. In contrast, offloading computationally intensive operations to an untrusted GPU greatly improves efficiency, but sacrifices the ability to verify computations performed outside the trusted boundary. Therefore, a natural direction is to seek an intermediate design that balances security and efficiency. Instead of requiring the protocol to rule out all possible adversarial behaviors, we tolerate certain adversarial behaviors under a controlled relaxation of the security guarantee. Such tolerance must satisfy three properties: • Negligible benefit from tolerated adversarial behaviors. Since a malicious trainer is primarily motivated by improving model utility, tolerated deviations should provide little or no utility advantage over honest execution. If deviating from the prescribed protocol yields negligible benefit, then the trainer’s incentive to cheat is substantially reduced.

Security–efficiency tradeoff. The above discussion reveals an efficiency-security tradeoff in verifiable DP training. Executing the entire training procedure inside a TEE provides strong integrity of DP guarantees but incurs prohibitive computational overhead. In contrast, offloading computation to untrusted GPUs achieves practical efficiency but creates opportunities for malicious deviations that undermine the integrity of the claimed DP execution. This raises a central challenge: how can we retain the efficiency of GPU-accelerated training while still providing strong, verifiable guarantee that the declared DP-SGD procedure has been faithfully executed?

5

Sparse and Dense Deviations

• No significant harm to privacy protection. Tolerated attacks should not introduce significant additional privacy leakage. • Significant efficiency improvement. The relaxation should enable substantial efficiency gains relative to fully trusted training, making verifiable DP training practical in realistic deployments. The first two properties bound the incentive and potential harm of tolerated deviations, while the third ensures that the relaxation yields meaningful efficiency gains. Together, these conditions suggest that a favorable balance between security and efficiency can be achieved by tolerating certain lowimpact adversarial behaviors. To identify such behaviors, we perform a finer-grained analysis of malicious deviations. Intuitively, some deviations can significantly alter the training trajectory, providing substantial utility gains to the trainer and potentially increasing privacy leakage. Other deviations affect only a small number of training steps. For the utility-oriented attacks evaluated in this work, we observe that such sparse deviations have only limited influence on final model utility and no measurable additional membership leakage. A natural way to distinguish between these two cases is by the number of malicious iterations during training. Let M be the total number of iterations on which the trainer deviates from the prescribed protocol, and let M be the maximum

Security–Efficiency Balance

In this section, we present our solution for verifiable DP training, which aims to achieve both practical efficiency and strong security. Our high-level idea is to tolerate sparse malicious deviations, which for the utility-oriented attacks we evaluate provide little utility benefit and no measurable additional membership leakage, while efficiently detecting dense malicious deviations that could meaningfully compromise the declared DP training execution. The section proceeds as follows. First, we introduce a frequency-based classification of malicious deviations and distinguish between dense and sparse attack regimes (Section 5.1). We then analyze the impact of sparse deviations on model utility and privacy leakage (Section 5.2): for the utility-oriented attack family evaluated in this work, sufficiently small attack budgets provide lim5

number of malicious iterations that the protocol is willing to tolerate. We classify malicious deviations below:

eration, the adversary selects the look-ahead checkpoint along the non-private trajectory as the target. • Nearest-checkpoint attack. The adversary selects the nonprivate checkpoint closest to the TEE-held model. • Final-checkpoint attack. Throughout the training process, the adversary can always fix the target checkpoint to the final checkpoint of the non-private model.

• Dense malicious activities. If M ≥ M, we refer to the malicious deviations as dense. • Sparse malicious activities. If M < M, we refer to the malicious deviations as sparse. This distinction captures an important asymmetry between impact and detectability. Dense deviations can substantially influence the optimization trajectory and therefore may provide meaningful utility gains to the trainer. However, because they occur on many iterations, they are also easier to detect through probabilistic verification. In contrast, sparse deviations are more difficult to detect because they occur on only a small number of iterations and would require substantially more verification effort to catch reliably. As we show later for the utility-oriented attacks we evaluate, sparse deviations offer only limited utility improvement and no measurable additional membership leakage. These observations motivate using deviations below the empirical threshold M as the tolerated operating regime in our system evaluation. However, two challenges remain. First, how should M be chosen to distinguish sparse attacks from dense ones in practice? Second, what are the concrete utility benefits and privacy harms introduced by deviations below this threshold? Later, we answer these questions through a comprehensive analysis of how the number of malicious iterations in DP training affects utility gain and privacy loss.

5.2

Under any of the above strategies the adversary also chooses which M iterations to deviate on, and we grant it the strongest choice rather than a random one. Both schedules concentrate on u⋆ , the fraction of training at which one unit of deviation buys the most displacement at the end (u⋆ = 0.20 for RoBERTa, 0.10 for GPT-2): the final-checkpoint attack takes the M consecutive steps centred on u⋆ T , while the current- and nearest-checkpoint attacks take one step from each window of width T /M, at the point of that window nearest u⋆ T . We did not run a uniform-random control, so these numbers should be read as the attacker’s best placement rather than as an average over placements. Given a selected target checkpoint, the details of how the adversary constructs the manipulated gradient to move the current model toward the checkpoint are provided in Appendix A; the per-setting constants and the remaining configuration are listed in Appendix B.1 and Table 4. 5.2.2

After characterizing the malicious deviation strategies, we now examine how the number of malicious iterations, M, affects the utility gain achievable by an adversary. We use this attack procedure to evaluate the utility benefit obtainable under different numbers of malicious iterations, focusing on M ∈ {40, 50, 60} in this section. Additional experiments with a broader range of M values are reported in Section 8. We conduct experiments on both conventional learning tasks and LLM-oriented fine-tuning tasks. For conventional learning, we consider a 5-layer MLP on Purchase and WideResNet on CIFAR-10. For LLM-oriented fine-tuning, we fine-tune GPT2-medium on E2E and WebNLG and RoBERTa-large on QQP and MNLI.1 We evaluate malicious DP training and honest DP training under privacy budgets ε = 2.0, along with nonprivate training, whose checkpoints serve as the attacker’s reference trajectory. Figure 1 summarizes the utility impact of sparse malicious deviations under ε = 2.0. For each task and each value of M, we report the best performance achieved among the three attack strategies, thereby representing the strongest utility gain obtainable by the adversary in our evaluation. The results show that deviations with at most 50 malicious iterations lead to only minor utility improvements over honest DP training in the experimental setting, especially when compared with

Analysis of Sparse Malicious Deviations

Now, we further analyze the effect of M on model utility enhancement and privacy leakage through both theoretical analysis and empirical evaluation. 5.2.1

Utility Analysis

Malicious Deviation Strategies

We first characterize the malicious deviation strategies considered in our analysis. As discussed earlier, a malicious trainer can steer the final DP model toward the corresponding nonprivate model by manipulating gradients during a subset of training iterations. Fundamentally, this attack aims to deviate the DP training trajectory toward a non-DP trajectory. Since a training trajectory can be viewed as a sequence of checkpoints, the adversary may locally maintain checkpoints from non-private training. Then, at each malicious iteration of DP training, the adversary selects a target non-DP checkpoint and manipulates the submitted gradient to move the TEEheld model toward that checkpoint. Below, we study three natural and straightforward strategies for selecting this target checkpoint. • Current-checkpoint attack. The most straightforward strategy is to select the checkpoint corresponding to the current training step. Specifically, at the t-th training it-

1We use LoRA for fine-tuning, a common parameter-efficient approach

for reducing the overhead of DP training.

6

Accuracy (%)

90

Purchase (MLP)

CIFAR-10 (WRN)

90.30

92.38

84 78 72

73.3674.2374.32 71.91

90

92

M = 50

M = 60

QQP (RoBERTa)

Non-private

MNLI (RoBERTa)

90.79

90

90

75

88

88

60

85.35 85.56 86 85.18 85.39

85.87 85.89 86 85.61 85.79

53.68 53.27 53.48 53.24

51

90.39

WebNLG (GPT-2)

45 42

E2E (GPT-2)

49.39

48

BLEU (%)

M = 40

DP

39.2039.3339.51 39 38.53

66.77

66 64 62.3062.2162.22

62 61.44

Figure 1: Utility comparison under sparse malicious deviations for ε = 2.0. The two RoBERTa panels use RoBERTa-large and the two GPT-2 panels GPT-2-medium. For each task, M = 40, 50, 60 report the best result among Current, Nearest, and Final attacks. Purchase, CIFAR-10, and QQP use accuracy, MNLI uses averaged matched/mismatched accuracy, E2E and WebNLG use BLEU, with the decoding configuration selected per bar on the validation split, so the bars within a panel do not share one configuration. The remaining metrics the two scorers emit are reported in Appendix C. the much larger utility gap between honest DP training and non-DP training. More precisely, when M = 50, all tasks except Purchase improve by less than one point over the DP baseline; even on Purchase, where the improvement is the most noticeable, the gain is only about 2.3 points. For CIFAR10 the deviations produce no consistent direction at all: the three budgets move accuracy by +0.20, −0.24 and −0.21 points, less than a quarter of a point either way. We therefore set M = 50 as the sparse–dense deviation threshold; this corresponds to 0.33%–0.89% of the training steps on the LLM workloads and 1% on CIFAR-10 (Table 4).2 Among the strategies evaluated here, none substantially closes the utility gap between honest DP and non-private training in this regime. We do not claim that the Current, Nearest, and Final strategies exhaust all possible utility-improving attacks: M = 50 is an empirical operating point supported by the strongest attacks in the evaluated family, and the protocol-level analyses of Sections 6 and 7.1 are stated independently of this family. Connection to bounded adversarial gradient perturbations. This assumption is consistent with a broader principle in optimization: when adversarial gradient perturbations are bounded in magnitude or frequency, their effect on the final model is also bounded. Recent theoretical work [32] shows that, for convex and smooth objectives, well-bounded gradient perturbations do not cause the learning process to deviate significantly; we discuss this connection in Appendix D.

scribed protocol on M iterations and, in the worst case, each deviating iteration may submit an arbitrary data-dependent vector of norm at most C, while the TEE still applies the prescribed Gaussian noise. If the original training procedure is (ε, δ)-DP under the add/remove-one neighboring relation, then, under the standard small-sampling-rate Gaussianaccounting approximation, the resulting procedure is approximately (ε′ , δ)-DP, where r ε′ M ≈ 1 + (4N 2 − 1). (1) ε T Does this theorem really imply a significant privacy loss? Theorem 5.1 implies a worst-case amplification of the privacy p loss by a factor of approximately 2N M/T when the malicious term dominates, which can indeed reach hundreds in practical settings. However, this seemingly large factor arises from an extremely pessimistic attack that is quite different from the utility-oriented deviations considered in our threat model. Specifically, the 4N 2 term allows the submitted vector on neighboring datasets to move between opposite points of the clipping ball, giving sensitivity up to 2C, and allows this worst-case dependence to concentrate repeatedly on the same target user. Repeating such highly targeted updates may maximize the privacy leakage of one particular user, but it provides little reason to expect a corresponding improvement in the overall utility of the trained model. As a result, a rational utility-oriented adversary has little incentive to conduct such attacks. In contrast, utility-oriented deviations such as the common ones discussed in Section 5.2.1 rely on broader, globally useful training signals rather than repeatedly concentrating the updates on a single user’s information. For illustration, suppose a manipulated gradient averages κ clipped user contributions in such a way that its per-user sensitivity is at most 2C/κ. Under the same approximation, s   M 4N 2 ε′ ≈ 1+ −1 . (2) ε T κ2

Theorem 5.1 (Privacy of DP Training with Incomplete Verification). Consider a T -iteration DP training procedure over a dataset of size N. Suppose the trainer deviates from the pre2 The attack budget M counts manipulated iterations; the full-deviation count Mfull of Section 7.1 counts iterations that a check would reject. For this construction the two coincide unless the honest aggregate nearly equals the forged one: the forged aggregate lies on the clipping-ball boundary, so an attacked iteration falls inside the tolerance region only if ∥gh ∥2 ≥ (1 − ρamb )C ≈ 0.99C, whereas the largest honest clipped average we measured was 0.65C on the GPT-2 tasks and 0.38C on the RoBERTa tasks. We treat the attacked iterations as full deviations on this basis, without re-verifying each one against the trusted reference.

7

Table 1: MIA performances on Purchase and CIFAR-10. DP stands for the DP-trained model, and M = 40, 50, 60 stands for the malicious deviated models. IMIA and SHAPOOL are two MIA methods, and TPR and AUC are corresponding metrics.

Thus, when κ is on the same order as N and M/T is small, the resulting privacy inflation can be close to one. This calculation is illustrative: dependence on many users alone does not imply the 1/κ sensitivity reduction; the latter requires the manipulated signal to average their contributions with correspondingly bounded per-user influence. To empirically validate this theoretical intuition, we next evaluate membership inference attacks and find that, for the utility-oriented sparse deviations evaluated (M ≤ 60), there is no measurable additional membership leakage beyond the variation of the honest DP baselines.

ε

Setting

IMIA: [email protected]% FPR Current Nearest

SHAPOOL: AUC

Final

Current Nearest Final

Purchase

Empirical study. To quantify the actual privacy leakage introduced by sparse malicious deviations, we evaluate two membership inference attacks: IMIA [12] and SHAPOOL [7], following the primary metric used in each attack’s original evaluation protocol. The detailed attack configurations are provided in Appendix G. As shown in Table 1, across both datasets Purchase and CIFAR-10, sparse deviations with M = 40-60 yield MIA performance very close to the corresponding honest DP baselines and substantially below that of non-private training. For SHAPOOL, the results are particularly stable: sparse malicious deviations achieve AUC values identical to, or within 0.01 of, those of the corresponding honest DP baselines, indicating no measurable additional membership leakage. For IMIA, the results exhibit greater variance. Nevertheless, the score values for both honest DP training and training with sparse deviations remain far below those of the corresponding non-private models. Interestingly, increasing either ε or M does not consistently increase the measured leakage. This non-monotonic behavior suggests that the stochastic variation inherent in DP training and MIA evaluation can be comparable to, or even larger than, the additional effect introduced by sparse malicious deviations. These results do not rule out sparse attacks designed to maximize privacy leakage: Theorem 5.1 characterizes a substantially worse case in which malicious iterations repeatedly expose one target user, whereas the attacks evaluated here use broad training signals.

Non-Private

–

2.0

DP M = 40 M = 50 M = 60

0.12 0.08 0.09 0.14

0.43 0.12 0.10 0.07 0.10

0.12 0.10 0.08 0.12

0.52 0.52 0.52 0.52

0.62 0.52 0.52 0.52 0.52

0.52 0.52 0.52 0.52

4.0

DP M = 40 M = 50 M = 60

0.08 0.10 0.09 0.10

0.08 0.09 0.08 0.15

0.08 0.08 0.10 0.11

0.52 0.52 0.52 0.52

0.52 0.52 0.52 0.52

0.52 0.52 0.52 0.52

Non-Private

–

2.0

DP M = 40 M = 50 M = 60

0.09 0.13 0.13 0.13

0.09 0.12 0.12 0.14

0.09 0.10 0.08 0.13

0.51 0.51 0.51 0.51

0.51 0.51 0.51 0.51

0.51 0.51 0.51 0.51

4.0

DP M = 40 M = 50 M = 60

0.08 0.07 0.09 0.11

0.08 0.10 0.07 0.09

0.08 0.08 0.09 0.14

0.50 0.50 0.50 0.50

0.50 0.50 0.50 0.50

0.50 0.50 0.51 0.50

CIFAR-10

6

1.14

0.62

Handling GPU–TEE Discrepancy

6.1 Honest Discrepancy and the Resulting Attack Surface Even when the GPU follows the protocol exactly, the clipped aggregate gradient it submits differs from the TEE’s own recomputation on the same batch and the same trusted state. The two sides execute different kernels (vendor GPU libraries versus CPU BLAS), reduce sums in different orders, fuse operations differently, and round at different points; per-example clipping can then amplify a rounding-level difference in a single example’s norm into a visible change of its clip factor. The discrepancy is therefore an intrinsic property of heterogeneous execution rather than a symptom of misbehavior, and a verifier that demanded bit-exact agreement would abort every honest run. Two empirical properties of this discrepancy shape our design. Let zt32 = ∥ḡtGPU − ḡt32 ∥2 /C denote the normalized distance between the GPU submission and the TEE’s FP32 recomputation. First, on the overwhelming majority of steps zt32 is tiny. Second, the distribution can have a heavy tail: rare, ill-conditioned steps—typically those containing examples whose per-example norm sits at the clipping boundary— produce discrepancies an order of magnitude larger than the body of the distribution. A single fixed tolerance is caught between two failure modes. Set near the body, it aborts honest runs on the tail steps; set near the tail, it hands a dishonest GPU a large per-step allowance on every step. Whatever tolerance the verifier grants for honest discrep-

Overall, for the utility-oriented sparse attacks studied here, we observe limited utility improvement and no measurable additional membership leakage. These findings motivate tolerating a small number of deviations from an incentive perspective; they are not a universal guarantee for arbitrary sparse strategies, and the protocol analysis that follows classifies arbitrary submitted gradients by their trusted verification outcome rather than relying on these attack constructions. The remaining task is therefore to design an efficient protocol that probabilistically detects repeated full deviations and accounts for the deviations admitted through numerical tolerance. 8

ancy is available to a dishonest GPU: a submission g̃t with ∥g̃t − ḡt32 ∥2 /C ≤ τ is indistinguishable from an honest one on that step. Probabilistic checking already accounts for the fact that unchecked steps are not examined at all—this is the source of the 1 − (1 − p)Mfull detection guarantee of Section 7.1. Numerical tolerance adds a second, subtler channel: even a checked step admits a deviation of up to τ, and such deviations, being individually below the tolerance, would never be flagged. Left unmetered, they accumulate without bound over a long run, so the adversary could steer the model through many small, undetectable pushes rather than a few large ones. Our design principle is therefore that no accepted deviation is free: every tolerated deviation observed on a checked step is charged to a bounded ledger, and hidden Bernoulli checking converts these sampled ledger charges into a high-probability bound on the cumulative deviation that can pass through the numerical-tolerance channels over the full run.

6.2

Algorithm 1: TEE Verification : GPU submission ḡtGPU ; verification probability p; thresholds τabs , τnum ; limits Ksub , Kamb ; clipping norm C; denominator B. GPU commits ḡtGPU ; if ∥ḡtGPU ∥2 > C(1 + 10−4 ) then return A BORT;

Input

TEE secretly samples Vt ∼ Bernoulli(p); if Vt = 0 then return ACCEPT; TEE recomputes ḡt32 and sets zt32 ← ∥ḡtGPU − ḡt32 ∥2 /C; if zt32 ≤ τabs then Zsub ← Zsub + zt32 ; if Zsub > Ksub then return A BORT; else return ACCEPT; B TEE recomputes ḡt64 and sets nt ← 2C ∥ḡtGPU − ḡt64 ∥2 ; if nt > τnum then return A BORT;

Two-Stage Verification

Rather than a single tolerance, the verifier uses an adaptive trusted reference: TEE FP32 recomputation is the reference by default, and steps whose FP32 discrepancy is unusually large are escalated to a TEE FP64 recomputation. Every submission, checked or not, first passes a structural invariant: the honest average of clipped per-example gradients satisfies ∥ḡtGPU ∥2 ≤ C by the triangle inequality, so the TEE rejects any submission with ∥ḡtGPU ∥2 > C(1 + 10−4 ), where the slack covers only FP32 norm-computation rounding. This check costs O(d), is applied on every step, and caps the gross deviation any single submission can carry. A checked step then follows one of three paths.

Samb ← Samb + 1; if Samb > Kamb then return A BORT; else return ACCEPT;

is made about why a step escalated or how many examples contributed to the discrepancy. Hard rejection. If zt64 > ρamb , the submission is inconsistent with the prescribed computation under either reference and the verifier aborts.

Body path. If zt32 ≤ τabs , FP32 serves as the trusted reference. The verifier does not simply accept: it charges the observed discrepancy to a cumulative ledger, Zsub ← Zsub + zt32 , and aborts once Zsub > Ksub . Thus τabs is a routing threshold rather than a free tolerance: an adversary that repeatedly hides just below τabs on checked steps exhausts Ksub after a bounded number of such steps, while honest runs, whose body discrepancies are far smaller than τabs , consume only a small fraction of the budget.

Algorithm 1 summarizes the procedure. Two remarks are in order. First, deviation is always measured relative to the trusted reference the TEE actually produces, so the GPU’s own numerical error never translates into attacker-controlled radius beyond what τabs and ρamb explicitly grant; conversely, the honest spectrum is a property of the specific GPU–TEE software and hardware pair, and the parameters are frozen per calibrated pair. Second, for CIFAR-10 the honest FP32 spectrum has no heavy tail, so the FP64 layer is unnecessary: escalated steps are instead accepted only under the samemetric hard cap zt32 ≤ ρill = 2τabs and counted by the same Kamb counter (Table 7).

Ambiguity path. If zt32 > τabs , the TEE recomputes the prescribed aggregate gradient in FP64, denoted ḡt64 , and evaluates zt64 = ∥ḡtGPU − ḡt64 ∥2 /C, equivalently nt = (B/2) zt64 in the accountant’s units. The step is accepted only if nt ≤ τnum , i.e. zt64 ≤ ρamb := 2τnum /B. The FP64 reference resolves the honest ill-conditioned cases, whose FP32 discrepancy was large only because FP32 rounding was amplified; but passing the FP64 check certifies plausibility, not honesty—an adversary may deliberately trigger escalation and hide within ρamb . We therefore treat every accepted fallback as residual numerical ambiguity and meter it separately with a counter: Samb ← Samb + 1, aborting once Samb > Kamb . No assumption

6.3

Calibration and False Abort

The four parameters play distinct roles: τabs controls routing, ρamb controls FP64 acceptance, Ksub meters cumulative body discrepancy, and Kamb meters accepted fallback events. The verifier parameters are obtained through a staged calibration process. We first use a broad numerical-discrepancy 9

campaign to characterize the body/tail structure of heterogeneous GPU–TEE execution and establish the two-stage calibration procedure. Before certified deployment, we apply this procedure to honest pilot trajectories on the target GPU–TEE pair, consolidate the resulting parameters at the model-family level, and freeze the complete verifier configuration for training. The verifier is calibrated for a concrete GPU–TEE numerical environment. Changes to the GPU family or trusted reference stack require recalibration; changes to the GPU-side training stack require revalidation and trigger recalibration if the resulting honest discrepancy spectrum falls outside the calibrated envelope. The verifier architecture, security accounting, and calibration procedure remain generic. The calibration trajectories are not identical in setup to the other experiments of this paper: they differ from the steeringattack experiments of Section 5.2 and the overhead runs of Section 8 in the weight-decay grouping (uniform versus the standard optimizer grouping; notes on Table 7, Appendix A), from the overhead runs additionally in the GPU-side software stack and the run length (Section 8), and from the deployed protocol in the audit law under which the numerical data were recorded; Appendix E documents each difference and its effect. The honest false-abort probability qFA is the probability, over the verifier’s coins alone, that an honest execution of a given trajectory is aborted; it is a property of that trajectory, not a prediction for future runs. It is evaluated under the deployed Bernoulli verification law as a tail bound on the sampled body charge, a binomial term for escalations, and (1 − p)Hr for hard rejections (Appendix E, Equation (8)). For a fully observed honest trajectory it is evaluated conditionally on the realized numerical sequence and requires no stationarity assumption; for trajectories observed only through a systematic every-tenth-step trace, the required trajectory statistics are estimated under an explicit representativeness assumption. Because Ksub and Kamb are absolute budgets while the honest audited charge grows with the number of steps, the reported values apply to runs of the calibrated length (Table 4; 5,000 steps for CIFAR-10); longer runs require rescaling Ksub and Kamb , with the corresponding change in Gextra . The design target is qFA ≤ 10−3 for every family. The family-level values in Table 7 are evaluated on the deployment-pair calibration trajectories themselves; they are therefore in-sample calibration checks, not estimates of the false-abort probability of any run. A false-abort estimate proper is available for one deployment trajectory only: after parameter freezing we preregistered a qqp-large trajectory with a fresh seed and replayed it with a full numerical census, obtaining an upper bound of 4 × 10−100 ; no held-out trajectory exists for the GPT-2 and CIFAR-10 families, so no false-abort estimate is reported for them. A hypothetical 20% inflation of the calibrated trajectory statistics keeps every family at or near this target, the thinnest margin being the GPT-2 body charge (Appendix E, which also gives the evidence tier of each reported quantity). The frozen verifier parameters, analytic security

budgets, in-sample calibration checks, and the single held-out false-abort estimate are listed in Table 7.

6.4

Security Budget of Numerical Tolerance

Numerical tolerance introduces a residual attack surface even on steps that would pass verification. Our ledgers meter such deviations whenever the corresponding step is checked, while hidden Bernoulli sampling allows us to bound the total deviation that can pass through these tolerance channels over the entire run. Let Gtol denote the cumulative normalized deviation routed through the body and ambiguity channels. We derive a highprobability security budget Gextra = Gsub + Gamb ,

(3)

such that Pr[Gtol > Gextra and still accepts] ≤ βextra := βsub + βamb . (4) For the normal FP32 path, Gsub follows from a Freedman– Bernstein bound for a predictable adversary that must commit before the current hidden Bernoulli coin is drawn. For the ambiguity path, Gamb follows from the exact binomial sampling law and the Kamb counter. Detailed derivations are given in Appendix E.2. At the primary operating point p = 0.1, the resulting numerical-tolerance security pairs (Gextra , βextra ) are (5.48, 0.05) for RoBERTa, (9.50, 0.05) for GPT-2, and (2.51, 0.05) for CIFAR-10 (Table 7), where the first coordinate is the cumulative normalized-deviation budget and the second the corresponding security failure probability. Importantly, Equation (4) is an analytic statement conditional on the frozen protocol parameters and the stated hidden-coin assumptions. Calibration determines the parameter values and hence the size of Gextra ; once they are frozen, the bound follows from the protocol and hidden Bernoulli sampling alone and does not assume that the adversarial discrepancy sequence follows the calibrated honest distribution. Appendix E.3 further gives an empirical same-state steeringscale interpretation of this normalized budget; it is used only to interpret the effect of numerical tolerance on the full-step detection term (Section 7.1), and the formal guarantee remains Equation (4).

7

Protocol Design

In this section, we design a verifiable training protocol that probabilistically detects repeated full deviations while accounting separately for deviations admitted through numerical tolerance. The core idea of our design is probabilistic verification. Instead of checking every training iteration, the TEE recomputes, or partially recomputes, gradients only on 10

a small randomly selected subset of iterations. Since gradient recomputation inside the TEE is much slower than GPU execution, such checks must remain infrequent in order to keep the system efficient. However, the checking schedule is hidden from the GPU worker, so the adversary does not know in advance which iterations will be audited. Consequently, an adversary that deviates frequently faces a correspondingly high probability of being detected. We now describe the basic protocol. Initialization. Before training begins, the TEE initializes and stores the trusted training state, including the model parameters, optimizer state, privacy parameters, and other global metadata. The GPU stores the training dataset and a local copy of the trainable parameters for efficient computation. Both parties share the public training specification, including the model architecture, loss function, optimizer, trainable parameter set, batch-selection procedure, and update rule. Training. At the t-th iteration, the TEE sends the GPU the current trainable model state and the information needed to identify the current batch. The GPU fetches the batch, performs the forward and backward passes, and returns the gradients of the trainable parameters to the TEE. The TEE then uses hidden randomness to decide whether this iteration should be spot-checked.3 If selected, the TEE recomputes, or partially recomputes, the expected gradient behavior on the same batch under the trusted state and compares it with the GPU output within an appropriate tolerance. If the discrepancy exceeds the tolerance, the TEE reports a deviation and terminates training. In addition, when persistent sealed storage and rollback protection are available, the TEE records a failure flag in its sealed state, so that subsequent attestations from the same TEE instance should not be certified. If the check passes, the TEE continues with the trusted DP-SGD operations, including DP-noise generation, privacy/accountingstate maintenance, optimizer-state evolution, and parameter update. In the optimized implementation of Section 7.2, perexample clipping and aggregation are performed on the GPU and probabilistically verified by trusted recomputation. The updated trainable state is then sent back to the GPU for the next iteration. Output. After training completes and every queued verification job has finished, the TEE releases the final model and attaches a TEE-generated certificate indicating that all audited iterations were consistent with the declared training protocol and that the numerical ledgers remained within their limits. The overall protocols are summarized in Algorithm 2 and 3. Overall, our design separates efficiency-critical computation from trust-critical computation. The GPU performs the

Algorithm 2: GPU-side Gradient Computation and Local Update Input : Training dataset D ; model f (·; θ) and loss ℓ; trainable parameters T ; clipping norm C; optimizer rule OptStep; batch seed sbatch ; initial (0) weights θT ; initial optimizer state ω(0) ; number of iterations T ; aggregation denominator B. (0) GPU: store D and receive (θT , ω(0) , sbatch ) from the TEE; (0) (0) b (0) ← ω(0) ; GPU: set b θT ← θT and ω

for t ← 0 to T − 1 do It ← BatchIdx(sbatch ,t); GPU: load Bt ← D [It ]; foreach (xi , yi ) ∈ Bt do (t) g ← ∇θ ℓ( f (xi ; b θ ), yi ); t,i

T

ḡt,i ← gt,i min{1,C/∥gt,i ∥2 }; ḡt ← B1 ∑(xi ,yi )∈Bt ḡt,i ; GPU → TEE: commit and send (t, It , ḡt ); GPU: wait for the TEE to release snoise,t ; e gt ← ḡt + B1 Noise(snoise,t ; σC);   (t+1) (t+1) (t) (t) b b ,e (b θT , ω ) ← OptStep b θT , ω gt ;

expensive training-side operations, while the TEE retains control over verification and privacy enforcement. In this way, the protocol substantially reduces the TEE-side burden compared with fully trusted training, while making repeated full deviations risky and explicitly accounting for manipulation admitted through numerical tolerance. In the following sections, we analyze the security guarantees of the protocol and introduce several optimizations that further reduce verification overhead without compromising security. Algorithm 3 is the blocking reference execution. The deployed hungry-updating implementation keeps the same certificate-time accept/reject semantics but reorders the runtime: once the submission is committed and the coin is sampled, a selected check is enqueued for background recomputation while the TEE proceeds with the provisional trusted update and seed release; the state remains provisional until every deferred job has completed successfully (Section 7.2).

7.1 Security Analysis & Incentive Engineering Accounting for numerical tolerance. An adversary may combine full deviations with deviations that remain inside the numerical-tolerance region; we account for the two components separately, each established on its own terms. First, deviations that a check would reject—submissions outside the tolerated region on their step—are full deviations. For a fixed set F of them,

3 The coin for step t is drawn from TEE-internal entropy after the stept submission has been received. The measurement harness used for the experiments of Section 8 instead audited a hidden uniformly random subset of ⌈pT ⌉ steps per run, fixing each run’s verification workload; Appendix E states what this affects.

Pdet (F ) = 1 − ∏ (1 − pt ). t∈F

11

(5)

Untrusted GPU

1

model state θt

CPU TEE (trusted)

next iteration

+ batch info

GPU

2 Forward / Backward

Recompute gt on Bt

compute gt on batch Bt

compare within tolerance

batch Bt

training data D

audit, w.p. p 3

gradient gt

mismatch → abort & report deviation

Verification pass 5 DP-SGD update

4 Hidden coin

skip

Clip

ct ~ Bern(p)

1−p

‖g‖ ≤ C

hidden from GPU

Add DP noise

Update model θt → θt+1

after T iterations Final model θT + TEE certificate audit, w.p. p

Iterations 1

T

Figure 2: Overview of our basic verifiable DP training protocol. The untrusted GPU performs the expensive forward and backward computation (steps 1–3), while the TEE uses a hidden coin Vt ∼ Bernoulli(p) to decide whether to audit the returned gradient by recomputation (step 4), and then executes the privacy-critical DP-SGD update (step 5). After T iterations, the TEE releases the final model with an attestation certificate. The figure shows the logical baseline; in the deployed implementation (Section 7.2, Algorithms 2–3) clipping and aggregation run on the GPU and are recomputed by the TEE only on audited steps. Since each iteration is audited independently with probability p and the schedule is hidden from the GPU, repeated full deviations are detected with increasing probability while the GPU retains its efficiency advantage. Under the uniform policy pt = p used throughout this work, the same argument applies to the first Mfull predictable fulldeviation opportunities generated by a history-adaptive adversary: each opportunity is determined before its hidden verification coin is drawn,4 so all Mfull opportunities are missed with probability (1 − p)Mfull and Pdet = 1 − (1 − p)Mfull . Second, deviations within the tolerated region are never individually flagged, but Section 6.4 bounds their total: the cumulative normalized deviation accepted through the tolerance channels over an entire run exceeds Gextra without the run being aborted with probability at most βextra = βsub + βamb . The fulldeviation bound and the pair (Gextra , βextra ) together constitute the formal security guarantee. They are stated in different units, require no empirical conversion between them, and both forms of deviation may coexist in the same execution. The uniform-p result and the pair hold for history-adaptive strategies satisfying the commit-before-current-coin condition; the nonuniform value-aware extension of Appendix E.4 is stated for a fixed set of full deviations.

(rejected-class), sub-threshold body, or ambiguity-path deviation. The uniform-p full-deviation bound and (Gextra , βextra ) therefore account for arbitrary submitted gradients in the verifier’s normalized-deviation metric, including strategies that adapt their direction, magnitude, or timing to the previous protocol history, subject to the commit-before-current-coin condition. What the formal accounting does not determine is how a given normalized deviation translates into optimization progress or model utility. Such an effect may depend strongly on the optimizer state, attack direction, and training phase. Appendix E.3 provides only an empirical calibration of this conversion over the tested attack states and radii. Attacks through channels other than the submitted aggregate gradient remain outside the present analysis. The formal quantities above do not depend on the attack experiments. The empirical operating point M = 50 (Section 5.2) and the steering-scale conversion of Appendix E.3 are measured on the attack-experiment trajectories, whose setup differs from that of the calibration trajectories (Section 6). Interpreting these empirical findings together with the security budgets instantiated for the deployment pair therefore assumes that their qualitative conclusions transfer across these trajectory families. We treat this as an empirical transfer assumption, not as part of the formal security guarantee. More generally, whenever an empirical quantity measured on one trajectory family is used elsewhere in the paper to interpret results instantiated on another trajectory family, the same transfer assumption should be understood unless stated otherwise.

Scope of the guarantee. Every committed submission that passes the structural check falls, according to the verdict a check would return, into exactly one of three classes: full 4 The coin is drawn only after the submission arrives, so no current-step decision exists at commitment time. What the GPU can observe is whether the previous step triggered extra TEE work (a timing channel); under independent coins this reveals nothing about later steps. Our analysis abstracts away timing and other side channels that could expose the current decision before commitment; a deployment can harden this by padding the latency-critical response path or by deferring verification dispatch to an asynchronous queue after commitment.

12

it does not include βextra , which is reported separately, and it interprets the cost of numerical tolerance only for attacks whose steering behaviour is represented by the calibration of Appendix E.3, not as a worst-case guarantee over arbitrary attack objectives. With Gextra = 0 and Asteer = Mfull it reduces to 1 − (1 − p)Mfull . For illustration, with Asteer = 40, p = 0.1, and the conservative Gextra = 10, the term decreases from 98.52% to 95.76%; under our calibrated budgets (Gextra ≤ 9.5) the change is smaller. Attack value may vary across training steps, but this timing information is also available to the verifier, which can allocate more of the same expected checking budget to the more sensitive phases so that higher-value attack steps also carry higher detection risk (Appendix E.4). All budgets and experiments in this paper use the uniform setting pt = p.

Algorithm 3: TEE-side Verification, Noise Release, and Update Input : Training dataset D ; model f (·; θ) and loss ℓ; trainable parameters T ; DP parameters (C, σ); optimizer rule OptStep; verification probability p; number of iterations T ; aggregation denominator B. (T ) Output : Final trusted weights θT . (0)

TEE: initialize trusted weights θT and optimizer state ω(0) ; TEE: choose master seeds sbatch and snoise ; (0)

TEE → GPU: send (θT , ω(0) , sbatch ); for t ← 0 to T − 1 do TEE: receive committed (t, It , ḡtGPU ) from the GPU; TEE: check It = BatchIdx(sbatch ,t); draw Vt ∼ Bernoulli(p) secretly; if Vt = 1 then TEE: load Bt ← D [It ]; foreach (xi , yi ) ∈ Bt do TEE ← ∇ ℓ( f (x ; θ(t) ), y ); gt,i i i θT TEE ← gTEE min{1,C/∥gTEE ∥ }; ḡt,i 2 t,i t,i

7.2

To make the split TEE-GPU design practical, we optimize both the communication path and the execution pipeline, which can substantially reduce the end-to-end overhead of verifiable DP training. Communication-efficient split execution. A straightforward implementation of the basic protocol would send all perexample gradients to the TEE and perform clipping inside the trusted environment. For a batch of size B and model dimension d, this requires O(Bd) GPU-to-TEE communication per iteration, which can dominate the runtime for modern models. We instead perform clipping on the GPU. The GPU computes the per-example gradients, clips each gradient, averages and submits only the resulting clipped-and-averaged gradient ḡtGPU to the TEE. The TEE retains responsibility for the privacy-critical randomness and optimizer state, while probabilistic verification checks whether the GPU-submitted aggregate is consistent with the prescribed computation. This reduces the GPU-to-TEE communication from O(Bd) to O(d) per iteration. We also avoid sending the updated model and optimizer state back to the GPU after every step. After the GPU commits ḡtGPU , the TEE releases the random seed used to generate the trusted DP noise for that step. The GPU reconstructs the same noise locally and applies the same optimizer update, thereby maintaining a synchronized copy of the model and optimizer state. The TEE-to-GPU communication is therefore reduced from O(d) to the size of a random seed. Releasing the seed only after the gradient commitment is essential: the GPU cannot adapt its submitted gradient to the realized DP noise. Thus, the common training path exchanges only one aggregate gradient in the GPU-to-TEE direction and a short random seed in the reverse direction. The more expensive per-example recomputation is incurred only on the subset of iterations selected for verification. Hungry updating with deferred verification. Gradient verification is substantially more expensive than the ordinary

TEE ; ḡtTEE ← B1 ∑(xi ,yi )∈Bt ḡt,i TEE: perform numerical verification using ḡtTEE ; if verification aborts then return A BORT;

snoise,t ← Derive(snoise ,t); TEE → GPU: release snoise,t ; ξt ← B1 Noise(snoise,t ; σC); e gt ← ḡtGPU + ξt ;   (t+1) (t) (θT , ω(t+1) ) ← OptStep θT , ω(t) ,e gt ; (T )

return θT ;

Empirical step-equivalent interpretation. Appendix E.3 calibrates one normalized deviation unit against one fullpower steering update and finds a ratio close to one over the tested regime. Measure each accepted deviation by its norm in units of the clipping bound: a step on which the honest aggregate is replaced by a full-power forgery contributes one unit, and a tolerated deviation of normalized radius r contributes r units, which is what the calibration above establishes. Writing Asteer for the sum of these contributions over a run — an attack’s full-power-equivalent steering magnitude, real-valued and distinct from the step count Mfull , though an attack that forges at full power on Mfull steps and nowhere else has Asteer = Mfull — the tolerance channels contribute at most about Gextra units, leaving (Asteer − Gextra )+ units outside the budget: interp

Pdet (Asteer ) ≈ 1 − (1 − p)(Asteer −Gextra )+ .

Efficiency Optimizations

(6)

Equation (6) is an interpretation: it rests on the empirical calibration and on treating tolerated and full deviations additively, 13

(a) Blocking verification

trusted update. In particular, a checked iteration requires the TEE to recompute the prescribed aggregate gradient in FP32, and a step exhibiting a large FP32 discrepancy may additionally require FP64 adjudication. Performing these computations synchronously would stall the GPU whenever a verification is triggered and would largely eliminate the benefit of probabilistic checking. We therefore decouple model updating from gradient verification using hungry updating. Once the GPU commits the aggregate gradient, the TEE executes the latency-critical DP update without waiting for the corresponding gradient-honesty check to finish. If the iteration is selected for verification, the TEE records the state needed to reproduce that iteration and places a verification job into a background queue. Training can then continue while background workers independently recompute the checked iterations. Because verification is deferred, a checked submission may enter the provisional training state before its job completes; such progress is not certified. A run is accepted—the meaning of the term throughout the security analysis—only when every verification job generated during the run has completed successfully and the certificate is issued; any failed deferred check aborts the run. Each queued job executes the verification procedure described in Section 6. Multiple verification workers can serve the queue in parallel, allowing training to run ahead of verification when sufficient CPU resources are available. Figure 3 illustrates the resulting pipeline. With these optimizations, the normal training path incurs only lightweight communication and trusted updating, while expensive recomputation is parallelized over probabilistically selected verification steps. In particular, when the aggregate verification capacity keeps pace with the arrival of checking tasks, i.e., n ≥ pTTiter-veri , most of the verification overhead can be hidden iter behind the main training pipeline, leaving only a small queuedraining cost at the end of training. Otherwise, verification becomes the throughput bottleneck. We provide a detailed workload and efficiency analysis in Appendix F.

8

audited

GPU

2

3

GPU idle

4

5

verify g3

TEE verify

GPU idle

GPU

1

2

3

4

5

6

enqueue

6

verify g5

(recompute)

(recompute)

time

(b) Hungry updating (deferred verification) training runs ahead of verification

verification queue

worker 1 worker 2

verify g3

(recompute)

verify g5

(recompute)

time saved

time

Figure 3: Hungry updating with deferred verification. (a) Blocking verification stalls training whenever a checked iteration is recomputed by the TEE. (b) Hungry updating places checked iterations into a background queue served by parallel verification workers, allowing training to proceed while verification executes asynchronously.

and one for CIFAR-10 (Table 7). The resulting adversarial budgets at the primary operating point p=0.1 are Gextra ≤ 5.48 (RoBERTa), 9.50 (GPT-2), and 2.51 (CIFAR-10), each with βextra = 0.05. The false-abort probability is estimated on one held-out deployment trajectory only, the preregistered qqp-large full-census replay run after the parameters were frozen, which gives an upper bound of 4 × 10−100 against the design target 10−3 . The family-level tail values on the deployment-pair calibration trajectories, 2.8 × 10−9 (RoBERTa), 8.7 × 10−6 (GPT-2), and 1.1 × 10−4 (CIFAR-10) under their respective evidence models (Appendix E), confirm that the frozen parameters meet the target on seven of the eight calibration trajectories (the qqp-large calibration trace is excluded, see Appendix E); they are in-sample calibration checks, not false-abort estimates. Generalization to future trajectories remains empirical. The overhead runs reported here differ from the deployment-pair calibration trajectories in three respects: the GPU-side software stack (the deployed PyTorch and PEFT versions, whereas the RoBERTa-base and GPT-2 calibration trajectories were produced under an earlier build), the optimizer grouping (the standard HF grouping that exempts biases and LayerNorm weights from weight decay, whereas the calibration trajectories apply it uniformly; notes on Table 7), and the run length (one epoch under the ten-epoch schedule rather than full trajectories). Appendix E details the software-stack difference and revalidates the frozen configuration against the honest discrepancy spectrum of these runs. The additional verification rates in Table 2 characterize the efficiency–rate tradeoff only; the budgets are not transferred to other rates without recomputation (with the same counters, the CIFAR-10 pair becomes (5.05, 0.05) at p=0.05 and (12.66, 0.05) at p=0.02).

Experiments

We conduct experiments on multiple datasets and models to evaluate our system. The evaluated models and datasets follow Section 5.2 and the details are provided in Appendix B.

8.1

1

Experimental setup

Hardware. We instantiate the untrusted trainer with a single NVIDIA RTX PRO 6000 GPU. The TEE-side verifier runs in an AMD SEV-SNP-protected VM configured with 64 CPUs and 512GB of memory. The GPU and the TEE communicate over a vsock channel [31]. Verification workflow. Checks are dispatched asynchronously so that verification overlaps GPU training. The judge parameters are frozen per model family — one configuration for all GPT-2 tasks, one for all RoBERTa tasks, 14

Table 2: End-to-end overhead and runtime diagnostics at two verification rates, excluding initialization and prewarming. TIn-GPU is the one-epoch unverified in-GPU DP training time. TVDP is the one-epoch training time of our protocol. TIn-TEE represents the estimated one-epoch training time of the full in-TEE solution. CIFAR-10 rows use a 200-step window for TVDP against a 500-step In-GPU baseline (its checks are ∼20× costlier, hence the lower rates p ∈ {0.02, 0.05}); the VDP overhead is therefore computed per warm step, excluding the first compiled step on both sides (per warm step: In-GPU 2.515 s versus VDP 2.843 s at p=0.02 and 3.284 s at p=0.05; the CIFAR-10 In-TEE ratio is also taken per warm step, and the raw totals in the CIFAR-10 row are not directly comparable across columns). For the LLM operating rates p=0.10 and p=0.15 shown here, Mfull = 50 full deviations are detected with probability 1 − 0.950 = 99.5% and 1 − 0.8550 = 99.97%, respectively (Eq. (5)); the CIFAR-10 rows use the lower rates p ∈ {0.02, 0.05} and are to be read with Eq. (5) at those rates; in addition, for the LLM cells the cumulative deviation admitted through the numerical-tolerance channels is bounded by (Gextra , βextra ) = (9.5, 0.05) or better (Section 7.1; the budgets of Table 7 are computed at the primary operating point p=0.1 and are conservative at p=0.15). Model

Task

p

WRN16-4

CIFAR-10

0.02 0.05

WebNLG GPT-2 Small E2E WebNLG GPT-2 Medium E2E QQP RoBERTa Base MNLI QQP RoBERTa Large MNLI

0.10 0.15 0.10 0.15 0.10 0.15 0.10 0.15 0.10 0.15 0.10 0.15 0.10 0.15 0.10 0.15

TIn-GPU (s) TVDP (s) TIn-TEE (s) Sync (s) Drain (s) Backpr. (s) VDP Overhead In-TEE 1306.04 109.12 119.08 277.84 301.79 582.96 631.72 1724.31 1873.26

743.29 801.80 115.24 122.01 127.29 140.20 290.53 310.39 315.11 344.75 667.27 723.43 726.14 797.76 1911.31 2176.47 2125.79 2439.39

25160.00 1001.03 1095.97 2298.14 2548.66 5003.06 5022.67 15590.00 15863.42

203.98 234.23

0.00 5.47

0.00 509.44

+13.0% +30.6%

5.11 12.04 6.32 15.32

0.00 2.28 3.13 4.01

0.40 7.08 0.93 10.23

+5.6% +11.8% +6.9% +17.7%

8.21 27.26 10.52 31.48

0.00 5.75 5.85 10.14

0.71 20.80 1.50 22.52

+4.6% +11.7% +4.4% +14.2%

47.09 93.34 53.34 117.26

0.36 0.74 0.84 1.47

7.97 48.00 20.56 86.86

+14.5% +24.1% +14.9% +26.3%

109.57 361.03 173.04 477.86

0.97 8.97 4.28 10.58

61.53 344.72 136.99 481.74

+10.8% +26.2% +13.5% +30.2%

20.0× 9.2× 9.2× 8.3× 8.4× 8.6× 8.0× 9.0× 8.5×

audits exactly ⌈pT ⌉ steps per run, i.e., the protocol’s mean checking workload; the reported times therefore exclude the run-to-run variation pthat the Bernoulli check count (relative standard deviation (1 − p)/(pT ), 6–13% on the LLM cells and 31–49% on the CIFAR-10 windows) and the placement of checked steps would add.

Benchmarks and baselines. We measure system overhead on both traditional learning tasks (CIFAR-10) and LLM-oriented tasks (E2E, WebNLG, MNLI, QQP), using the same family of models as in Section 5.2. We omit Purchase because its training schedule is too short to yield meaningful epoch-level timing. Our floor-level baseline is unverified DP training on a single GPU, which incurs no communication or checking cost. We also report a TEE-only baseline that runs the entire DP training procedure inside the CPU-TEE, representing the naive alternative. It is implemented as a single-machine CPU DP-SGD trainer inside the same SEV-SNP guest, reusing the exact DP recipe of the verified runs (per-example clipping, seeded noise, optimizer, and schedule) with the batch sharded across the guest’s vCPUs at the measured optimal shape (32 workers × 2 threads). Because a full in-TEE epoch takes hours to days, we time five optimizer steps after a warm-up step and extrapolate linearly to one epoch. For CIFAR-10 the in-TEE time is derived from the measured in-guest cost of the verifier’s full-batch recomputation (the same computation as one training step) on all 64 vCPUs. The measurement harness

Metrics. We report wall-clock timing for a single training epoch, excluding initialization and worker prewarming. TIn-GPU denotes the runtime of unverified in-GPU DP training on the same GPU, while TVDP denotes the end-to-end runtime of our protocol, including the final completion of all verification jobs generated during the epoch. We report the following diagnostics: • Sync, the cumulative GPU-side waiting time after submitting an update and before receiving the corresponding noise seed from the TEE. It includes gradient serialization, GPU–TEE communication, deserialization, and TEE response latency, and is therefore an upper bound on pure communication cost. 15

+14.5%

+24.1%

150

0

178 137 93

100 50

+40.5%

Sync (GPU wait) Drain Backpressure

200

cost relative to K = 1

exposed time (s)

250

48

47 0

8

p = 0.10

3

1

p = 0.15

p = 0.20

exposed time (s)

1000

+10.8%

+26.2%

+57.5%

200 0

110 1

62

p = 0.10

345

1.0

+3%

K=1 S = 62

K=2 S = 31

K=4 S = 15

17

9

p = 0.15

p = 0.20

8.2

(b) QQP / RoBERTa-large

• Drain, the time required to complete verification jobs that remain pending after GPU training finishes. This captures the verification work that cannot be hidden behind training.

End-to-End Overhead. We report the results in Table 2. Running the entire training loop inside the TEE is prohibitively slow, incurring 8.0×–9.2× overhead on the LLM tasks and 20× on CIFAR-10, whereas our protocol stays below 15% on every LLM task at the protocol operating point p=0.10, and below 31% even at the elevated rate p=0.15. At p=0.10 the drain column is at most a few seconds on all eight tasks, indicating that almost all verification work completes by the time GPU training ends. Verification cost can nevertheless be exposed during training through synchronization and backpressure, particularly on the larger RoBERTa workloads; on the remaining tasks the residual overhead is almost entirely per-step synchronization. For example, on QQP with RoBERTa-base the end-to-end overhead is 14.5%, of which 47.09 s out of the 84.31 s of added wall-clock time is Sync; the RoBERTa tasks pay a visibly larger synchronization share than the GPT-2 tasks. Raising the rate to p=0.15 moves the large models toward the TEE-paced regime: on MNLI with RoBERTa-large, Sync (477.86 s) and backpressure (481.74 s) grow together — the two overlap, as the training loop increasingly waits on check throughput rather than on transport — and the overhead reaches 30.2%. The verification rate p thus acts as a direct dial between auditing intensity and exposed cost, with the entire measured range remaining substantially cheaper than in-TEE training (6.3×–8.7× across the LLM cells).

• Backpr., the time training is blocked because the number of in-flight verification jobs reaches the protocol limit. A nonzero value indicates that verification throughput becomes a bottleneck. Since Backpr. is measured on the TEE side while Sync is measured on the GPU side, the two may overlap and should not be added directly.

• End-to-End Overhead, the end-to-end slowdown relative to unverified in-GPU training.

Table 3: Model accuracy of WideResNet on CIFAR-10 under different privacy budgets with M = 120 malicious iterations. ε

Setting

Nearest

∞

Non-Private

2.0

DP M = 120

53.48% 54.20%

53.48% 52.81%

53.48% 53.83%

4.0

DP M = 120

65.80% 65.09%

65.80% 64.83%

65.80% 65.81%

8.0

DP M = 120

72.23% 83.21%

72.23% 72.99%

72.23% 85.38%

Efficiency Evaluation

We first evaluate the efficiency of our verification protocol. The main questions are: (i) how much overhead our protocol adds over ordinary single-GPU DP training, (ii) how much cost is avoided compared with the naive TEE-only design, and (iii) where the remaining cost comes from.

Figure 4: Exposed cost components vs. verification rate p on QQP. Bold annotations give the end-to-end overhead.

Current

+18%

1.1

Figure 5: Verifier shape scan on the deployment pair (QQP / RoBERTa-large, p=0.20; identical audited-step set across configurations; pool size fixed at K ·S ≈ 62). All costs are relative to K=1.

600 361

+29%

933

888

800

400

1.2

0.9

(a) QQP / RoBERTa-base

per-check compute (idle bench) per-check latency (deployed, p50) end-to-end wall clock

1.3

Final

92.38%

Overhead regimes. Figure 4 dissects the exposed cost as the verification rate grows, for the same task at two model sizes. Two regimes are visible. For RoBERTa-base (Fig. 4a), per-step synchronization dominates at every rate: checks are cheap enough (166 core-seconds each) that the verifier pool 16

keeps pace with training, backpressure stays below Sync even at p=0.20, and the overhead is essentially the price of the perstep seed round-trip. For RoBERTa-large (Fig. 4b), whose checks cost 3.4× more, the system crosses into a throughputbound regime: backpressure grows from about half of Sync at p=0.10 (62 s vs. 110 s) to overtaking it at p=0.20 (933 s vs. 888 s), and much of the Sync growth in this regime is itself backpressure propagating to the GPU-side seed wait — the training loop increasingly waits on check throughput rather than on transport. Drain stays within a few seconds in every cell: backpressure throttles training in place of letting unfinished checks spill past the epoch boundary, so verification adds at most a small tail after training completes. Larger models therefore reach the throughput-bound regime at lower verification rates, which is precisely the regime where enlarging the verifier’s worker pool (rather than reducing p) recovers the overhead.

hardware-backed isolation and remote attestation, making them useful for protecting or verifying ML computation on untrusted accelerators. Slalom [40] combines a CPU-side TEE with an untrusted GPU to verify outsourced computation. Later systems use partitioning or obfuscation to protect sensitive model components while outsourcing most computation to the GPU [17, 22, 25, 38, 44, 48]. For training integrity, TrustFL [47] and GINN [6] use TEE-based sampled verification of untrusted GPU computation, with GINN further combining gradient clipping and asynchronous verification. TrustFL+ [23] addresses GPU–TEE floating-point non-determinism, while AFTUNE [20] uses TEE-based spot checking for outsourced fine-tuning. In contrast, we focus on DP training under incomplete TEE verification and explicitly quantify the residual adversarial freedom introduced by unchecked steps and numerical tolerance. GPU TEEs and confidential GPU computing. Another line of work explores ML integrity through GPU TEEs [18, 37, 41, 43]. However, GPU-TEE support remains limited in deployed infrastructures [44], and many training clusters still rely on legacy GPUs. Our work is complementary: instead of requiring trusted GPUs, we target existing CPU-TEE and legacy-GPU environments.

Choosing the verifier shape. The verifier uses K concurrent checks, each split into S shards, with K·S fixed by the available vCPUs. Figure 5 shows that increasing K provides little throughput benefit but increases per-check latency: at K=4, each check receives fewer workers, increasing its p50 latency by 29% and the end-to-end runtime by 18%. A larger K also keeps more verification states in flight, increasing memory pressure and queueing overhead. We therefore use small K (K ≤ 2) and choose S per task. This result also shows that verifier shape should be selected from end-to-end measurements.

10

In this paper, we presented a CPU-TEE-based framework for verifiable DP training on untrusted GPUs. The protocol provides analytic accounting for arbitrary submitted aggregate gradients through probabilistic detection of full deviations and a high-probability numerical-tolerance budget. Separately, for the utility-oriented forged-gradient attacks evaluated in our experiments, sparse attack budgets provide limited utility improvement and show no measurable additional membership leakage; these empirical findings motivate the operating point used by our system rather than constituting a worst-case guarantee over all possible sparse attacks. Our evaluation shows that the design adds modest overhead over standard GPUbased DP training while avoiding the substantially higher cost of running the full training procedure inside a CPU-side TEE.

Additional experiments on CIFAR-10. We further examine how many malicious steps are required for an effective attack on CIFAR-10. From Table 3, we can see even 120 malicious steps offer little gain at ε = 2 or 4, and become effective only at ε = 8. In contrast, RoBERTa on MNLI already shows small but consistent gains with only 40–60 malicious steps at ε = 2. Hence, effective attacks on CNNs require a substantially larger M, which directly benefits verification efficiency.

9

Conclusion

Related Work

Verifiable differential privacy. Verifiable differential privacy certifies the correct execution of claimed differentially private computations. Narayan et al. [26] first introduced this concept and showed how verifiable computation can certify differentially private data analysis. Biswas et al. [8] applied zeroknowledge proofs to DP counting queries, and Noisette [30] certifies DP noise sampling for discrete and continuous mechanisms. More recently, Confidential-DPProof [34] and VeriDP [2] use customized zero-knowledge proofs to verify differentially private model training. These approaches provide strong cryptographic assurance without trusted hardware, but incur substantial proof-generation overhead for realistic training workloads. In contrast, we explore a TEE-based design that trades complete verification for efficient probabilistic auditing.

References [1] Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pages 308–318, 2016. [2] Behzad Abdolmaleki, Amir R Asadi, Vahid R Asadi, Stefan Köpsell, Bhavish Mohee, Nahid Roustaeifar, and Maryam Zarezadeh. Veridp: Verifiable differentially private training. Proceedings on Privacy Enhancing Technologies, 2026.

TEE-assisted verifiable machine learning. TEEs provide 17

[3] John M Abowd. The us census bureau adopts differential privacy. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 2867–2867, 2018.

[14] Cynthia Dwork and Aaron Roth. The algorithmic foundations of differential privacy. Foundations and trends® in theoretical computer science, 9(3-4):211–487, 2014.

[4] Erdem Aktas, Cfir Cohen, Josh Eads, James Forshaw, and Felix Wilhelm. Intel trust domain extensions (tdx) security review. Google security review, 2023.

[15] Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. Model inversion attacks that exploit confidence information and basic countermeasures. In Proceedings of the 22nd ACM SIGSAC conference on computer and communications security, pages 1322–1333, 2015.

[5] Emily Apsey, Phil Rogers, Michael O’Connor, and Rob Nertney. Confidential computing on nvidia h100 gpus for secure and trustworthy ai, august 2023. URL https://developer. nvidia. com/blog/confidentialcomputing-on-h100-gpus-for-secure-and-trustworthyai/. Accessed, pages 7–17, 2024.

[16] Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. The webnlg challenge: Generating text from rdf data. In Proceedings of the 10th international conference on natural language generation, pages 124–133, 2017. [17] Jiahui Hou, Huiqi Liu, Yunxin Liu, Yu Wang, Peng-Jun Wan, and Xiang-Yang Li. Model protection: Real-time privacy-preserving inference service for model privacy at the edge. IEEE Trans. Dependable Secur. Comput., 19(6):4270–4284, 2021.

[6] Aref Asvadishirehjini, Murat Kantarcioglu, and Bradley Malin. Ginn: fast gpu-tee based integrity for neural network training. In Proceedings of the Twelfth ACM Conference on Data and Application Security and Privacy, pages 4–15, 2022.

[18] Weizhe Hua, Muhammad Umar, Zhiru Zhang, and G. Edward Suh. Guardnn: Secure DNN accelerator for privacy-preserving deep learning. CoRR, abs/2008.11632, 2020.

[7] Li Bai, Qingqing Ye, Xinwei Zhang, Sen Zhang, Zi Liang, Jianliang Xu, and Haibo Hu. Toward efficient inference attacks: Shadow model sharing via mixtureof-experts. Advances in Neural Information Processing Systems, 38:138757–138785, 2026. [8] Ari Biswas and Graham Cormode. Verifiable differential privacy. arXiv preprint arXiv:2208.09011, 2022.

[19] Matthew Jagielski, Jonathan Ullman, and Alina Oprea. Auditing differentially private machine learning: How private is private sgd? Advances in Neural Information Processing Systems, 33:22205–22216, 2020.

[9] Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer. Membership inference attacks from first principles. In 2022 IEEE symposium on security and privacy (SP), pages 1897–1914. IEEE, 2022.

[20] Heng Jin, Chaoyu Zhang, Hexuan Yu, Shanghao Shi, Ning Zhang, Y Thomas Hou, and Wenjing Lou. Trusting what you cannot see: Auditable fine-tuning and inference for proprietary ai. arXiv preprint arXiv:2603.07466, 2026.

[10] Rishav Chourasia, Ergute Bao, Uzair Javaid, and Xiaokui Xiao. Auditing apple’s differentialprivacy. framework: Implementation bugs, misconfigurations, and practical risks. arXiv preprint arXiv:2605.21378, 2026.

[21] Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. [22] Ding Li, Ziqi Zhang, Mengyu Yao, Yifeng Cai, Yao Guo, and Xiangqun Chen. Teeslice: Protecting sensitive neural network models in trusted execution environments when attackers have pre-trained models. ACM Transactions on Software Engineering and Methodology, 34(6):1–49, 2025.

[11] Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar, Pasin Manurangsi, Amer Sinha, and Chiyuan Zhang. Scalable DP-SGD: Shuffling vs. Poisson subsampling. In Advances in Neural Information Processing Systems (NeurIPS), 2024.

[23] Cheng Lyu, Xiaoli Zhang, Jiaqing Cheng, Wenmao Liu, Xiaohu Ye, Ke Xu, Qi Li, and Xu-Cheng Yin. Towards efficient and reliable training assurance of untrusted federated learning participants under hardware non-determinism. IEEE Transactions on Dependable and Secure Computing, 2026.

[12] Yuntao Du, Yuetian Chen, Hanshen Xiao, Bruno Ribeiro, and Ninghui Li. Imitative membership inference attack. arXiv preprint arXiv:2509.06796, 2025. [13] Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. Calibrating noise to sensitivity in private data analysis. In Theory of cryptography conference, pages 265–284. Springer, 2006.

[24] Ilya Mironov. Rényi differential privacy. In 2017 IEEE 30th computer security foundations symposium (CSF), pages 263–275. IEEE, 2017. 18

[25] Fan Mo, Ali Shahin Shamsabadi, Kleomenis Katevas, Soteris Demetriou, Ilias Leontiadis, Andrea Cavallaro, and Hamed Haddadi. Darknetz: towards model privacy at the edge using trusted execution environments. In Eyal de Lara, Iqbal Mohomed, Jason Nieh, and Elizabeth M. Belding, editors, MobiSys ’20: The 18th Annual International Conference on Mobile Systems, Applications, and Services, Toronto, Ontario, Canada, June 1519, 2020, pages 161–174, 2020.

[36] Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), pages 3–18, 2017. doi: 10.1109/SP.2017.41. [37] Supraja Sridhara, Andrin Bertschi, Benedict Schlüter, Mark Kuhne, Fabio Aliberti, and Shweta Shinde. ACAI: protecting accelerator execution with arm confidential computing architecture. In Davide Balzarotti and Wenyuan Xu, editors, 33rd USENIX Security Symposium, USENIX Security 2024, Philadelphia, PA, USA, August 14-16, 2024. USENIX Association, 2024.

[26] Arjun Narayan, Ariel Feldman, Antonis Papadimitriou, and Andreas Haeberlen. Verifiable differential privacy. In Proceedings of the Tenth European Conference on Computer Systems, pages 1–14, 2015.

[38] Zhichuang Sun, Ruimin Sun, Changming Liu, Amrita Roy Chowdhury, Long Lu, and Somesh Jha. ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural Networks . In IEEE Symposium on Security and Privacy (SP), pages 1596–1612. IEEE Computer Society, 2023.

[27] Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. The e2e dataset: New challenges for end-to-end generation. In Proceedings of the 18th annual SIGdial meeting on discourse and dialogue, pages 201–206, 2017. [28] Santa Clara NVIDIA. Nvidia blackwell architecture technical brief, 2024.

[39] Jun Tang, Aleksandra Korolova, Xiaolong Bai, Xueqiang Wang, and Xiaofeng Wang. Privacy loss in apple’s implementation of differential privacy on macos 10.12. arXiv preprint arXiv:1709.02753, 2017.

[29] OECD. Emerging privacy enhancing technologies: Current regulatory and policy approaches. Technical report, Organisation for Economic Co-operation and Development, 2023.

[40] Florian Tramer and Dan Boneh. Slalom: Fast, verifiable and private execution of neural networks in trusted hardware. In International Conference on Learning Representations, 2019. URL: https://openreview. net/forum?id=rJVorjCcKQ.

[30] Qi Pang, Radhika Garg, Ziling Liu, Hanshen Xiao, Virginia Smith, Wenting Zheng, and Xiao Wang. Noisette: Certifying differential privacy mechanisms efficiently. Cryptology ePrint Archive, 2026. [31] Rusty Russell. virtio: towards a de-facto standard for virtual I/O devices. ACM SIGOPS Operating Systems Review, 2008.

[41] Stavros Volos, Kapil Vaswani, and Rodrigo Bruno. Graviton: Trusted execution environments on gpus. In Andrea C. Arpaci-Dusseau and Geoff Voelker, editors, 13th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2018, Carlsbad, CA, USA, October 8-10, 2018, pages 681–696. USENIX Association, 2018.

[32] Nawapon Sangsiri and Yufei Tao. Distributed Learning with Adversarial Gradient Perturbations. https://www. cse.cuhk.edu.hk/~taoyf/paper/ijcai26.pdf, 2026. Accessed 2026-06-07.

[42] Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Glue: A multitask benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP, pages 353–355, 2018.

[33] AMD Sev-Snp. Strengthening vm isolation with integrity protection and more. White Paper, January, 53(2020):1450–1465, 2020. [34] Ali Shahin Shamsabadi, Gefei Tan, Tudor Cebere, Aurélien Bellet, Hamed Haddadi, Nicolas Papernot, Xiao Wang, and Adrian Weller. Confidential-dpproof: Confidential proof of differentially private training. In International Conference on Learning Representations, volume 2024, pages 4029–4044, 2024.

[43] Chenxu Wang, Yunjie Deng, Zhenyu Ning, Kevin Leach, Jin Li, Shoumeng Yan, Zhengyu He, Jiannong Cao, and Fengwei Zhang. Building a lightweight trusted execution environment for arm gpus. IEEE Transactions on Dependable and Secure Computing, 2023.

[35] Ali Shahin Shamsabadi and Nicolas Papernot. How to deploy machine learning with differential privacy, 2023. URL: https://www.nist.gov/blogs/ cybersecurity-insights/how-deploy-machinelearning-differential-privacy.

[44] Pengli Wang, Bingyou Dong, Yifeng Cai, Zheng Zhang, Junlin Liu, Huanran Xue, Ye Wu, Yao Zhang, and Ziqi Zhang. Game of arrows: On the (in-)security of weight 19

One-step AdamW map. Given a candidate forged gradient g, define the noise-free one-step AdamW map Φt : Rd → Rd by

obfuscation for on-device tee-shielded llm partition algorithms. In 34th USENIX Security Symposium (USENIX Security 25), pages 279–298, 2025.

m(g) = β1 mt−1 + (1 − β1 )g,

[45] Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, 2018.

v(g) = β2 vt−1 + (1 − β2 )g ⊙ g, b (g) = m

Φt (g) = θt − ηt

[46] Lukas Wutschitz, Huseyin A. Inan, and Andre Manoel. dp-transformers: Training transformer models with differential privacy. https://www.microsoft.com/enus/research/project/dp-transformers, August 2022. Code: https://github.com/microsoft/dptransformers.

b v(g) =

v(g) , 1 − βt2 !

b (g) m p + λθt b v(g) + ε

.

The current model parameters and optimizer states are fixed throughout the inner optimization. Noise-free surrogate objective. the one-step surrogate

The attacker minimizes

ft (g) = D(Φt (g)) ,

[47] Xiaoli Zhang, Fengting Li, Zeyu Zhang, Qi Li, Cong Wang, and Jianping Wu. Enabling execution assurance of federated learning at untrusted participants. In IEEE INFOCOM 2020-IEEE Conference on Computer Communications, pages 1877–1886. IEEE, 2020.

subject to g ∈ BC . Since our experiments fine-tune LoRA parameters and, for classification tasks, an additional classification head, we define the steering distance as

[48] Zheng Zhang, Na Wang, Ziqi Zhang, Yao Zhang, Tianyi Zhang, Jianwei Liu, and Ye Wu. Groupcover: A secure, efficient and scalable inference framework for on-device model protection based on tees. In Fortyfirst International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL: https://openreview.net/ forum?id=4mU6LNMaIu.

D(θ) =

∑

LoRA pairs j

2

s j B j A j − s j B⋆j A⋆j F +

∑

∥θi − θ⋆i ∥22 ,

unpaired i

where s j = α j /r j is the LoRA scaling. For LoRA parameters, the distance is measured on the merged update ∆W j = s j B j A j rather than directly on the factors (A j , B j ). This removes the ambiguity caused by the nonunique LoRA factorization: for any invertible G ∈ Rr j ×r j , the reparameterization (A⋆j , B⋆j ) 7→ (GA⋆j , B⋆j G−1 ) leaves the merged update unchanged,

[49] Jianwei Zhu, Hang Yin, Peng Deng, Aline Almeida, and Shunfan Zhou. Confidential computing on nvidia hopper gpus: a performance benchmark study. arXiv preprint arXiv:2409.03992, 2024.

A

m(g) , 1 − βt1

(B⋆j G−1 )(GA⋆j ) = B⋆j A⋆j , so a distance taken on the factors would charge for the choice of G while a distance taken on the product does not. Consequently, the attack objective depends only on the target merged update rather than on a particular choice of target factors.

Optimizer-Aware Forged-Gradient Attack

Setup and threat model. Let θt ∈ Rd denote the trainable parameters at step t. The optimizer is AdamW with first- and second-moment states (mt−1 , vt−1 ), hyperparameters β1 , β2 ∈ [0, 1), learning rate ηt , numerical constant ε, and decoupled weight decay λ (applied per parameter group as configured by the trainer; in the attack experiments the standard HF grouping is used, i.e. λ = 0 for biases and LayerNorm weights). The adversary may forge the aggregated gradient g consumed by the optimizer, subject to the clipping constraint

A.1

Solving the Inner Problem

Linearized initialization. We initialize the attack using the constrained minimizer of the first-order expansion of ft around g = 0. Let wt := ∇g ft (g) g=0 . Then the minimizer of the linearized objective over BC is

g ∈ BC := {x : ∥x∥2 ≤ C}.

g(0) = −C The realized DP noise is not observed when g is selected. 20

wt . ∥wt ∥2

(7)

1 − cos = 1.2 × 10−14 (QQP/roberta-base) and 5.7 × 10−14 (E2E/gpt-2), both at t = 0.6 T and for the seed alone — seven orders below the fp32 precision the deployed solver runs at. It is recorded because “read from the live optimizer state” would otherwise be read as exact.

The gradient in (7) is obtained by automatic differentiation through the complete surrogate D(Φt (g)). Equivalently, wt = J ⊤ 0 ∇θ D(Φt (0)) ,

J0 =

∂Φt , ∂g g=0

where J 0 is diagonal, so the product above is elementwise. In particular, the steering gradient is evaluated at Φt (0), rather than at θt , since the existing AdamW first moment and weight decay can move the parameters even when the newly submitted gradient is zero. Because (7) differentiates the same merged-space objective used by the attack, the initialization is independent of the particular factorization used to represent the target LoRA adapters.

B

Description of Used Datasets

We evaluate our method on datasets covering tabular classification, image classification, natural language understanding (NLU), and natural language generation (NLG) tasks. Purchase [36]. Purchase is a tabular classification dataset widely used in privacy and membership-inference studies. Each record represents a user’s purchase behavior, and the goal is to predict the corresponding purchase category. We use this dataset to evaluate DP training on multilayer perceptrons (MLPs), which serve as representative models for conventional tabular learning tasks. CIFAR-10 [21]. CIFAR-10 is a standard image classification benchmark consisting of natural images from ten object categories. We use CIFAR-10 to evaluate DP training on convolutional neural networks (CNNs), where the model is trained to classify each image into one of the ten classes. The model is a WideResNet-16-4 trained on 40,000 target training images, with 15,000 images reserved as shadow data for the membership-inference experiments and 5,000 as the test set; the honest DP-SGD runs and the attacked runs share the same recipe: ε = 2 (Table 3 additionally reports ε = 4 and 8), δ = 10−5 , 5,000 training steps, batch size 2,048 with microbatches of 512, learning rate 2.0, and clipping norm C = 1. MNLI [45]. The Multi-Genre Natural Language Inference (MNLI) dataset is a natural language understanding benchmark. Given a premise and a hypothesis, the task is to predict whether the hypothesis is entailed by, contradicts, or is neutral with respect to the premise. We use MNLI to evaluate DP fine-tuning for sentence-pair understanding tasks. QQP [42]. The Quora Question Pairs (QQP) dataset is another natural language understanding benchmark. Given a pair of questions, the task is to determine whether the two questions are semantically equivalent. We use QQP to evaluate DP fine-tuning on semantic matching and paraphrase detection tasks. E2E [27]. The E2E dataset is a natural language generation benchmark for data-to-text generation. Given a structured meaning representation, the task is to generate a fluent natural-language utterance that accurately describes the input attributes. We use E2E to evaluate DP fine-tuning for controlled text generation. WebNLG [16]. WebNLG is a data-to-text natural language generation dataset built from structured RDF triples. Given a set of subject–predicate–object triples describing entities and relations, the model must generate a coherent naturallanguage description that preserves the input facts. Compared

Projected gradient refinement. Starting from g(0) , we refine the forged gradient using projected gradient descent on the exact surrogate: ! ∇g ft (g(k) ) (k+1) (k) g = ΠBC g − γkC , ∥∇g ft (g(k) )∥2 where   C ΠBC (x) = x min 1, . ∥x∥2 We use backtracking line search: a candidate is accepted only if it is finite and decreases ft ; otherwise the step size is halved and retried. On acceptance the step size is doubled for the next iteration, capped at γ, so the search can recover a large step after a locally difficult region. The optimization terminates after at most K iterations, when no backtracking step decreases the objective, or when the change in g falls below tolerance τ. Noisy execution. Let g⋆ denote the forged gradient returned by the inner optimization. The actual optimizer consumes    1 C ⋆ e g = g min 1, ⋆ + ξ, ξ ∼ N 0, σ2C2 I . ∥g ∥2 B Thus the attack is optimized using the noise-free surrogate, while the realized AdamW update and optimizer states are formed from the noisy gradient e g. ηt is taken from the live parameter group before the scheduler has written the current step’s value, so the surrogate sees ηt−1 . It cancels from the solution almost entirely: ηt multiplies the whole Jacobian and so leaves any normalised direction unchanged, surviving only through the point Φt (0) at which ∇D is read. Nor does that residue grow late in training — under a linear decay ηt − ηt−1 = η0 /T is constant while Φt (0) → θt . Rebuilding the seed at both rates, each with ∇D re-evaluated at its own Φt (0), turns its direction by 21

with slot-based datasets such as E2E, WebNLG covers more diverse domains and relational structures, making it useful for evaluating DP fine-tuning on structured generation tasks.

B.1

No metric shows a substantial absolute gain. On E2E/gpt2medium, where the attacked runs do move every metric in the same direction as BLEU, the largest changes are +0.014 METEOR, +0.113 NIST, +0.081 CIDEr and +0.005 ROUGE-L — 0.7 to 4.1 percent of the metric’s own value, against nonprivate references that remain 0.04 to 0.33 away. On E2E/gpt2 the metrics do not even agree in sign: NIST and METEOR rise while BLEU, ROUGE-L and CIDEr fall, which is what one expects when the effect is smaller than the noise of selecting a decoding configuration on a 547-example validation split, and we therefore draw no conclusion from that setting. On both WebNLG settings every non-BLEU change other than TER is zero or a single unit in the last printed digit, and TER improves by 0.02 and 0.04; the evaluator prints those metrics to two decimals, so they cannot resolve an effect of this size in either direction. The reading consistent with all four settings is that sparse deviations do not produce a clear gain on any generation metric. Where the metrics are precise enough to resolve the effect at all, the absolute movement is small; where they are not, they are silent rather than supportive.

Attack Configuration

Table 4 lists the per-setting constants.5 The remaining choices are shared across all eight settings and are stated here. Budgets and strategies. The deviation budget is M ∈ {40, 50, 60} and each budget is run under all three target strategies (final, nearest, current), giving 3 × 3 × 8 = 72 attacked runs alongside the eight honest DP baselines and the eight non-private references. Which steps are deviated on. The adversary chooses the step set rather than drawing it at random, so the reported numbers are its best placement. Both schedules concentrate on u⋆ , the fraction of training at which one unit of deviation buys the most displacement at the end. The final-checkpoint attack, whose target is fixed, injects on the M consecutive steps centred on u⋆ T . The nearest- and current-checkpoint attacks instead take one step from each window of width T /M, at the point of that window closest to u⋆ T . We did not run a uniform-random control. Inner loop. The forged gradient is refined by projected gradient descent from the linearised seed: step size γ = 1 as a fraction of the clipping norm C, at most K = 100 iterations, up to ten halvings of the trial step per iteration, and a stopping tolerance of 10−6 on ∥g(k+1) −g(k) ∥2 . A candidate is accepted only if it is finite and decreases the surrogate. Non-private reference. For the classification tasks the target is the final non-private checkpoint. For E2E and WebNLG6 it is the non-private checkpoint with the best validation score, which is not the last one in three of the four settings; the column in Table 4 names the checkpoint used in each case.

D

Analogy to Learning with Bounded Adversarial Gradient Perturbations

We further discuss a related theoretical perspective from distributed learning with bounded adversarial gradient perturbations. This line of work considers a setting where a learner queries gradients from clients, but each returned gradient may be adversarially perturbed subject to a bounded-distance constraint. More concretely, for a client loss function ℓi , the returned vector u need not equal the true gradient ∇ℓi (w); it only needs to satisfy ∥u − ∇ℓi (w)∥2 ≤ ε.

C

Generation Metrics Beyond BLEU

Under convex and L-smooth objectives, the study shows that such adversarial perturbations do not make learning arbitrary. Instead, when an upper bound ∥w⋆ ∥2 ≤ R on the optimal solution is known, the achievable optimization error is controlled by the perturbation radius and problem parameters: the minimum unavoidable sub-optimality is on the order of εR, and an algorithm can guarantee a sub-optimality gap of order εR with finite query complexity. Although our setting is different, this result provides a useful analogy for understanding why sparse malicious deviations in our protocol provide limited utility benefit. In the deployed protocol, per-example clipping and aggregation are executed on the untrusted GPU, but every submitted aggregate is subject to the always-on norm invariant ∥ḡtGPU ∥2 ≤ C(1 + 10−4 ), while the prescribed honest clipped aggregate ḡtref has norm at most C. By the triangle inequality,

Figure 1 reports BLEU for E2E and WebNLG. Both scorers emit further metrics, and since a single metric can move for reasons unrelated to output quality we report all of them here. Table 5 gives, for each setting and each metric, the DP baseline, the most favourable of the nine attacked runs for that metric, and the non-private reference. Taking the best cell per metric rather than per run is deliberate: it asks whether any attacked run moves the metric at all, which is the more demanding question. 5We use the dp-transformers trainer [46]: fixed-size batches drawn from a per-epoch reshuffle with the last partial batch dropped, and RDP/PRV accounting at sampling rate B/N. Such means has been acknowledged as a common practice [11]. 6 On the generation tasks the per-example loss is the token mean over that example’s own unmasked positions and the backward scalar is the mean over examples, so each per-example gradient is the gradient of one record’s loss rather than of a batch-wide token mean.

∥ḡtGPU − ḡtref ∥2 ≤ ∥ḡtGPU ∥2 + ∥ḡtref ∥2 ≤ (2 + 10−4 )C. 22

Table 4: Parameters of the forged-gradient experiments. All eight settings share ε = 2, clipping norm C = 1, learning rate 4e − 4 with a linear schedule and no warm-up, weight decay 10−2 , AdamW (β1 , β2 ) = (0.9, 0.999), εAdam = 10−8 , 10 epochs, seed 1, and LoRA dropout 0 with every other dropout module zeroed. T is the number of optimizer steps (the sampler drops the last partial batch, so T = ⌊N/B⌋×epochs). σ is the noise multiplier the accountant returns at that (ε, δ, B/N, T ); the per-coordinate noise actually added is σC/B. u⋆ is the centre of the injection window as a fraction of training. The target is the non-private checkpoint the adversary steers towards under the final strategy. task

model

QQP QQP MNLI MNLI E2E E2E WebNLG WebNLG

roberta-base roberta-large roberta-base roberta-large gpt2 gpt2-medium gpt2 gpt2-medium

N

B

T

363,846 363,846 392,702 392,702 42,061 42,061 18,025 18,025

256 256 256 256 64 64 32 32

14,210 14,210 15,330 15,330 6,570 6,570 5,630 5,630

seq r α 128 128 128 128 128 128 256 256

8 8 8 8 4 4 4 4

Hence every accepted submission lies within (2 + 10−4 )C of the trusted aggregate. This provides a bounded-perturbation analogy, analogous to the bounded adversarial-gradient model above, where each adversarial reply may deviate from the true gradient only within a finite radius; it should not be interpreted as the formal security theorem of our non-convex stochastic training setting. This analogy supports the intuition that the norm invariant limits the adversary’s per-step influence. The adversary may still bias the optimization trajectory, but it cannot inject an unbounded update through a single gradient reply. Why the related result is more general in adversarial strength? The bounded-perturbation setting considered in the prior theoretical study is more general than our tolerated sparse-deviation regime in several important ways. First, it allows every gradient query to be adversarially perturbed, as long as the returned vector remains within the prescribed perturbation radius. In contrast, our relaxed security goal only tolerates sparse deviations: the adversary may deviate on a small fraction of iterations, while repeated full deviations are detected with high probability by our probabilistic checking mechanism. Second, the adversarial perturbation in that model is worst-case and can be chosen adaptively at each query, whereas in our protocol the submitted aggregate is subject to the always-on aggregate-norm invariant and to the trusted DP-SGD update rule, while GPU-side clipping and aggregation are probabilistically checked by trusted recomputation. Third, our training process additionally includes DP noise injected inside the TEE, which is not controlled by the trainer and further limits the trainer’s ability to precisely steer the final model. Therefore, the prior result should not be interpreted as a direct proof for our non-convex, stochastic DP training setting. Nevertheless, it gives a useful conceptual justification: even in a stronger setting where the adversary can perturb every gradient within a bounded radius, the resulting opti-

8 8 8 8 8 8 8 8

δ

σ

10−6 10−6 10−6 10−6 10−5 10−5 10−5 10−5

0.6462 0.6462 0.6407 0.6407 0.6679 0.6679 0.6827 0.6827

σC/B u⋆ / target 0.00252 0.00252 0.00250 0.00250 0.01044 0.01044 0.02133 0.02133

0.20 / checkpoint-14220 0.20 / checkpoint-14220 0.20 / checkpoint-15340 0.20 / checkpoint-15340 0.10 / checkpoint-3290 0.10 / checkpoint-5922 0.10 / checkpoint-5076 0.10 / checkpoint-5640

mization error remains controlled rather than arbitrary. Since our tolerated adversary is further restricted to sparse deviations, the cumulative benefit in our setting is intuitively even more limited. The related result therefore provides a useful bounded-perturbation analogy rather than a formal guarantee for our non-convex stochastic setting. In our protocol, clipping and the aggregate-norm invariant constrain the magnitude of each submitted update, while the limited utility impact of sparse deviations is established empirically for the attack family evaluated in Section 5.2.

E Calibration and Security Cost of Numerical Tolerance This appendix complements Section 6 with the precise definitions, the calibration rules, and the formal accounting behind the numerical-tolerance budget. For a checked step, define ∥ḡtGPU − ḡt32 ∥2 ∥ḡGPU − ḡt64 ∥2 , zt64 = t . C C The threshold τabs routes a step to FP64 adjudication. An escalated step is accepted only if zt32 =

2τnum . B Thus, ρamb is the security-relevant aggregate acceptance radius; τnum is only its implementation-specific normalization. zt64 ≤ ρamb ,

E.1

ρamb =

Honest Calibration and False Abort

For a fixed honest trajectory r, partition the steps into 32 Ur = {t : zt,r ≤ τabs }, 32 64 Ar = {t : zt,r > τabs , zt,r ≤ ρamb }, 32 64 Hr = {t : zt,r > τabs , zt,r > ρamb }.

23

Table 5: Every metric the official scorers emit, on the test split at ε = 2. “best attack” is the most favourable of the nine attacked runs for that metric, so different rows may be won by different (M, strategy) cells; ∆ is its absolute change from the DP baseline, in the metric’s own units. TER is lower-is-better, so a negative ∆ is an improvement there. The WebNLG evaluator prints every metric but BLEU to two decimals, which is the granularity of the changes observed on that task. At that granularity the best value is usually reached by several cells at once, and the last column then reports how many rather than naming one. setting

metric

DP

best attack

∆

non-private

E2E / gpt2

BLEU NIST METEOR ROUGE_L CIDEr

0.6044 7.8193 0.3833 0.6211 1.8375

0.5999 7.9538 0.3906 0.6156 1.8128

-0.0045 +0.1345 +0.0073 -0.0055 -0.0247

0.6495 8.2836 0.4219 0.6726 2.1242

M=40, current M=40, current M=60, final M=60, nearest M=60, current

E2E / gpt2-medium

BLEU NIST METEOR ROUGE_L CIDEr

0.6144 8.0637 0.4087 0.6563 1.9659

0.6230 8.1765 0.4224 0.6609 2.0464

+0.0086 +0.1128 +0.0137 +0.0046 +0.0805

0.6677 8.5079 0.4608 0.7015 2.3245

M=40, nearest M=40, final M=50, final M=50, nearest M=50, final

WebNLG / gpt2

BLEU BLEU_NLTK METEOR chrF++ TER

34.55 0.3300 0.3100 0.5300 0.6100

35.26 0.3400 0.3200 0.5400 0.5900

+0.71 +0.0100 +0.0100 +0.0100 -0.0200

46.82 0.4600 0.3800 0.6400 0.5000

M=50, final 9 of 9 cells tied 9 of 9 cells tied 4 of 9 cells tied 4 of 9 cells tied

WebNLG / gpt2-medium

BLEU BLEU_NLTK METEOR chrF++ TER

38.53 0.3800 0.3500 0.5900 0.5900

39.51 0.3800 0.3500 0.5900 0.5500

+0.98 +0.0000 +0.0000 +0.0000 -0.0400

49.39 0.5000 0.4000 0.6700 0.4300

M=60, final 9 of 9 cells tied 3 of 9 cells tied 2 of 9 cells tied 5 of 9 cells tied

Vt ∼ Bernoulli(p), and all security and false-abort probabilities reported in this paper are evaluated under this law. The numerical data used to select and validate the verifier parameters were collected through several measurement procedures: (i) the broad calibration campaign, 32 trajectories on a calibration GPU–TEE pair, carries z64 on every step and z32 on the harness’s uniformly random audits (rate 0.1) plus probed tail steps. (ii) The deployment-pair systematic traces, eight honest pilot trajectories on the target pair, carry every-tenthstep z32 and full-step FP64 records; seven of them (all but qqp-large) are also the trajectories on which the family-level tail values of Table 7 are evaluated. (iii) The qqp-large replay, a trajectory preregistered with a fresh seed after parameter freezing and recorded on every step, is the only held-out trajectory and hence the only false-abort estimate; and (iv) the overhead harness fixes a hidden uniformly random audit subset of size ⌈pT ⌉ in order to stabilize the verification workload during timing measurements. Because honest verification is passive (the audited recomputation does not alter the trajectory; replay fidelity was verified bitwise), the underlying trajectory statistics (Sr , Qr , Mr , Hr ) do not depend on the audit law; the observation procedure determines only how completely those statistics are measured. For the systematic traces of (ii) we treat the every-tenth-step sample as representative of the trajectory when forming the trajectory-level

Let Mr = |Ar |,

Hr = |Hr |,

and define the sampled body charge 32 Zr = ∑ Vt zt,r ,

Vt ∼ Bernoulli(p).

t∈Ur

Writing (r)

Fsub (x) = Pr[Zr ≤ x], and conditioned on the honest trajectory satisfying the always-on aggregate-norm invariant (observed on all calibration trajectories and the held-out replay), the exact trajectory-conditional false-abort probability under independent Bernoulli checking is (r)

Hr qFA r = 1 − Fsub (Ksub )FBin (Kamb ; Mr , p)(1 − p) .

best cell

(8)

Equation (8) is conditional on the realized honest trajectory and makes no stationarity assumption on numerical errors. Sampling laws in calibration and deployment. The deployed protocol uses independent hidden verification coins 24

up

up

up

upper estimates Sr , Qr , Mr (no period-ten structure in the numerical discrepancy); the resulting values are model-based estimates under this representativeness assumption rather than distribution-free confidence bounds, and on the census trajectory the estimates obtained from each of the ten offsets cover the exact statistics. The hard-rejection count Hr is exact on every trajectory and does not rely on the assumption. For CIFAR-10 the statistics come from the harness’s audited steps, a uniformly random subset for which the binomial count inversion is conservative. Table 6 summarizes the data sources.

discrepancy spectrum: over the 3770 checks judged in those runs, the ratio of the observed body-discrepancy means to those of the calibration traces lies in 0.81–1.23 for six of the eight tasks and below 0.81 for the remaining two (the conservative direction), with a single escalation, accepted by the FP64 check (nt = 0.91 ≤ τnum ), and no norm violation; this lies inside the margins of Table 7 and within the +20% sensitivity row of the evaluation paragraph, so we retain the frozen parameters for these runs. The equivalence of the TEE reference itself was tested once rather than assumed: recomputing recorded calibration steps of both families on the host reference environment and inside the SEV-SNP guest gave bitwiseidentical discrepancies, and recomputing one recorded step with the verifier’s shard grouping (micro-batches of 256, 8, and 4 examples; 16 and 1 threads) reproduced the recorded value bitwise, so the host-side calibration passes are taken as the deployed TEE computation. The FP64 references of the calibration stages were computed on a GPU, whereas the deployed verifier computes them on the TEE CPU; recomputing four recorded steps on the TEE CPU (two from the campaign and two from the deployment-pair traces, including the two largest FP64 discrepancies observed, nt ≈ 1.0) reproduced the GPU values to within 3 × 10−9 relative, an absolute difference in z64 below 10−12 , ten orders of magnitude below ρamb , which we treat as negligible. For CIFAR-10 the calibration and deployment runs share one GPU environment and one TEE environment; in a small-scale test on a single audited step, host and guest recomputations agreed bitwise, and recomputing that step with chunk sizes of 2048, 256, 33 (the verifier’s shard width), and 8 examples changed z by at most 10−7 relative (fp32 summation order; not bitwise), three orders of magnitude below τabs ; on this small-scale evidence we treat the effect as negligible.

Staged calibration. The verifier parameters are obtained through the calibration procedure developed from the broad numerical-discrepancy campaign and subsequently instantiated on the target GPU–TEE pair. The broad campaign characterizes the body/tail structure, fallback behavior, and crossrun variation needed to establish the calibration rules; its four runs per configuration share the data order and differ only in a shifted noise stream, so the cross-run variation it exhibits understates that of independent replicates, and no reported quantity relies on it beyond the +20% sensitivity row. Before certified deployment, we apply these rules to honest pilot trajectories on the target pair, consolidate the resulting values at the model-family level, and freeze the complete verifier configuration. A change in GPU family or in the trusted reference stack triggers recalibration; a change in the GPU-side training stack is revalidated against the calibrated discrepancy envelope and triggers recalibration if the observed honest spectrum falls outside it. Environment differences between the calibration stages and the deployed pair. Not all calibration trajectories were produced under the deployed GPU software stack. For the RoBERTa-large configurations the GPU side of both calibration stages ran under the deployed stack; for RoBERTa-base and the four GPT-2 configurations it ran under an earlier PyTorch build and a PEFT version whose LoRA initialisation differs, so those calibration trajectories belong to a different honest-trajectory family than the deployed runs. The TEE-side reference is unaffected, because every reference recomputation copies the trainable state from the GPU-side record. The fallback radius and the escalation-count inputs, which the procedure takes from the broad campaign, were checked on the deployed pair: the full-step FP64 census of the eight deployment-pair traces and of the held-out replay stays below ρamb (largest 7.81 × 10−3 against 1.185 × 10−2 for RoBERTa, 6.58 × 10−4 against 4.891 × 10−3 for GPT-2, so Hr = 0 exactly), and the escalation counts observed on the deployed pair (at most two on any every-tenth-step trace, 22 on the full-census replay) stay below the campaign-derived opportunity bounds and far below Kamb . The deployment runs of Section 8 likewise use a GPU-side software stack that differs from the deployment-pair calibration trajectories, so we revalidate the frozen configuration against their honest

In-sample calibration checks and the held-out estimate. The family-level values of Table 7 are evaluated on seven of the eight honest trajectories used during deployment-pair calibration (all but qqp-large; see the evaluation paragraph), the same trajectories on which the calibration procedure was instantiated. They confirm that the frozen parameters meet the design target on those trajectories, under the corresponding evidence model, and are therefore in-sample calibration checks. A false-abort estimate is available for exactly one deployment trajectory: after the parameters were frozen we preregistered a new qqp-large trajectory with a fresh seed, ran it on the target pair under the deployed software stack, and replayed it with full-step numerical recording; this held-out replay passed the verifier with a trajectory-conditional false-abort upper bound of 4.1 × 10−100 and retuned no parameter. Evaluating Equation (8). For a trajectory with a full census, Mr and Hr are read directly, and for the fully observed 32 } body sequence {zt,r t∈Ur the independence of the Bernoulli 25

Table 6: Provenance of the honest numerical data.

E.2

data source

observation pattern

role

broad calibration campaign (32 traj., calibration pair) deployment-pair traces (8 traj., target pair)

z64 every step; z32 on random audits + probes

reference characterization

every-tenth-step z32 ; fullstep FP64

qqp-large preregistered replay (target pair) CIFAR-10 traces (4 traj., target pair)

every step

deployment-pair calibration; in-sample check held-out evaluation deployment-pair calibration; in-sample check

z on harness audits (random subset)

Exhaustive classification of GPU submissions. For every committed GPU submission that satisfies the always-on structural checks, the trusted recomputation defines a counterfactual verification outcome independently of whether that step is actually sampled. The submission therefore belongs to exactly one of three disjoint classes: the FP32 body B , the accepted FP64 ambiguity region A , or the hard-rejection region H . The first two classes constitute the numerical-tolerance channels and are accounted for by (Gextra , βextra ) below. Every step in H is rejected whenever sampled and is therefore covered by the Bernoulli full-step detection guarantee: if at least Mfull predictable hard opportunities occur, the first Mfull of them are all missed with probability (1 − p)Mfull , again for any history-adaptive strategy, and violations of the alwayson structural invariants are rejected deterministically. This classification does not depend on how the GPU constructs its submitted gradient; a more elaborate forging strategy may change which class a step falls into, but not the classification itself. Throughout, “accepted” means that the run reaches certificate issuance after every deferred verification job has completed.

checking coins gives the Chernoff bound   32 (r) 1 − Fsub (Ksub ) ≤ inf exp(−λKsub ) ∏ 1 − p + p eλzt,r . λ>0

Security Accounting

t∈Ur

(9) When the body was recorded on a systematic sample (every tenth step), the realized moments are replaced by trajectorylevel upper estimates computed as nominal simultaneous 95% confidence bounds for a uniform random sample of the same size; these are model-based estimates under the assumption that the systematic sample is representative of the trajectory (see the sampling-laws paragraph above), not distributionfree confidence bounds. For CIFAR-10, where discrepancies are recorded only on the harness’s randomly audited steps, the honest escalation rate is instead bounded by the rule of three on the calibration audits, FBin (Kamb ; Mr , p) is replaced by the corresponding binomial bound over the ≈500 checked steps of the 5,000-step run at p=0.1, and the hard-cap term (1 − p)Hr is supported empirically only (no honest audited step exceeded ρill during calibration). Within the corresponding evidence model, each family-level value in Table 7 is at least the largest evaluated per-trajectory upper value within that family. The qqp-large calibration trace is excluded from the RoBERTa entry: under the systematic-sample evidence model the range term of its body estimate leaves the bound at about 0.2, far above the target, which is why that configuration was re-examined by the full-census replay. As a sensitivity check rather than a bound, inflating every calibrated up up up trajectory statistic (Sr , Qr , Mr , and the CIFAR-10 escalation rate) by 20% gives worst-case conditional values of 1.3 × 10−6 (RoBERTa), 4 × 10−4 to 1 × 10−3 (GPT-2, depending on whether the second-moment estimate is inflated linearly or quadratically), and 2.9 × 10−4 (CIFAR-10); a 30% inflation exceeds the 10−3 target for GPT-2. The thinnest marup gin is the GPT-2 body charge, pSr /Ksub = 0.73 on the worst trajectory (E2E, GPT-2-medium), against at most 0.43 for RoBERTa.

Normal FP32 path. Let at denote the adversarial normalized deviation from the TEE FP32 reference on a body-path step, with 0 ≤ at ≤ τabs . The adversary may choose at adaptively from the previous protocol history, but must commit before the hidden currentstep verification coin Vt is revealed. Define GT = ∑ at , t

ZT = ∑ Vt at . t

The cumulative check accepts only while ZT ≤ Ksub . Using the martingale difference Xt = (p −Vt )at and stopping when the cumulative deviation first reaches G, a Freedman–Bernstein bound gives Pr[GT ≥ G, ZT ≤ Ksub ] ≤ (pG − Ksub )2 exp − 2p(1 − p)τabs (G + τabs ) + 23 τabs (pG − Ksub )

! .

(10) We define Gsub by numerically inverting Equation (10) at failure probability βsub . Hence, cumulative body-path deviation exceeding Gsub evades this counter with probability at most βsub . 26

FP64 ambiguity path. at64 =

Table 7: Frozen verifier parameters per model family at the primary security operating point p = 0.1; the numericaltolerance guarantee is the pair (Gextra , βextra ) with βextra = βsub + βamb = 0.05 in every family. z-thresholds are in clipping-norm units (C=1); τnum is stated in the accountant’s amplified units (×B/2C) and scales with the batch size (0.1565 at B=64, 0.0783 at B=32). CIFAR-10 replaces the fp64 layer with the same-metric hard cap ρill = 2τabs . For the two false-abort rows see the notes below the table.

Every accepted fallback use satisfies

∥ḡtGPU − ḡt64 ∥2 ≤ ρamb . C

Consider the first M ambiguity opportunities generated by any predictable adversarial strategy, where whether the current submission constitutes such an opportunity is determined before the current verification coin is drawn. Because the verification coins are independent across steps and hidden until commitment, the audit indicators on these first M opportunities are i.i.d. Bernoulli(p), so the number of audited ones satisfies

RoBERTa

XM ∼ Binomial(M, p). Define Mβamb = max {M : FBin (Kamb ; M, p) > βamb } . Surviving beyond Mβamb opportunities requires the acceptedescalation counter to stay at or below Kamb on the first Mβamb + 1 of them, hence  Pr at least Mβamb + 1 ambiguity opportunities occur  and the counter stays ≤ Kamb ≤ βamb ,

audit rate p τabs ρamb / ρill

0.1 1.60 × 10−3 1.185 × 10−2

τnum Ksub Kamb (βsub , βamb ) aggregate-norm slack

1.5169 0.088 25 (.005, .045) 10−4 (1, 64) / (2, 32)

worker pool (K, S)

GPT-2

CIFAR-10

0.1 0.1 1.74 × 10−3 7.30 × 10−4 4.891 × 10−3 1.459 × 10−3 0.1565 / 0.0783 — 0.72 0.20 17 5 (.025, .025) (.025, .025) 10−4 10−4 (2, 32)

(62×1)

5.478 9.496 2.511 ≤ 2.8 × 10−9 in-sample check (target 10−3 ) (3 traj.) ≤ 8.7 × 10−6 ≤ 1.1 × 10−4 −100 ≤ 4 × 10 held-out qFA estimate (qqp-large replay) — —

Gextra

for every history-adaptive strategy. The corresponding cumulative allowance is Gamb = ρamb Mβamb . Notes on Table 7. The two false-abort rows are evaluated under the Bernoulli protocol from the statistics observed as described in the sampling-laws paragraph. The held-out row is a trajectory-conditional upper bound from the exact statistics of the full-census replay and is the only false-abort estimate in this paper. The in-sample row is evaluated on the deployment-pair calibration trajectories: the RoBERTa entry covers three of the four RoBERTa trajectories (the qqp-large calibration trace being excluded as explained in the evaluation paragraph), the RoBERTa and GPT-2 entries are model-based upper values under the representativeness assumption for the systematic every-tenth-step sample, and the CIFAR-10 entry uses the rate model over the ≈500 checked steps of the 5,000step run; they certify only that the frozen parameters meet the target on the calibration data. Weight-decay convention for the LLM families (CIFAR-10 uses no weight decay): the released verifier and GPU trainer apply weight decay uniformly to every trainable parameter on both sides, matching the broad calibration campaign and the deployment-pair calibration trajectories (single-group AdamW). The overhead runs reported here were executed with the standard HF optimizer grouping on both sides, which exempts biases and LayerNorm weights from decay; the grouping changes the training trajectory but not the verifier implementation or worker configuration. The steering-attack experiments of Section 5.2 and Appendix A also use the standard HF grouping.

CIFAR-10 specialization. For CIFAR-10, which does not use FP64 adjudication, the same accounting applies with the ambiguity region defined directly in the FP32 metric:

B = {t : zt32 ≤ τabs }, A = {t : τabs < zt32 ≤ ρill }, H = {t : zt32 > ρill }. Each accepted ambiguity opportunity contributes at most ρill , so Gamb = ρill Mβamb with the same Bernoulli-counter argument. Combining the two channels, Gextra = Gsub + Gamb .

(11)

If the cumulative normalized deviation Gtol routed through the body and ambiguity channels exceeds this total, then at least one of the two component allowances is exceeded. Hence, Pr[Gtol > Gextra and the run is accepted] ≤ βsub + βamb =: βextra . A submission with zt64 > ρamb does not belong to Gextra and is rejected whenever sampled; the choice of reference on escalated steps is discussed under Status of the Guarantee below. The parameter table is given below. 27

E.3

Empirical Steering-Scale Calibration

we therefore report the interpreted full-step detection term as approximate, while the deviation bound of Equation (11) holds irrespective of this calibration. Two limits bound what this measures. The sub-threshold and ambiguity branches are not independent evidence: they share one linearisation and their Rt agree to five decimals, so their agreement is arithmetic rather than corroboration. And the matched noise draw is what makes the ratio tight — it is the correct comparison for a single step, since the noise is common to both branches and cancels along the first-moment path, but under independent draws the same quantity varies by a factor of two to three across seeds.

This subsection is empirical. The formal guarantee is stated in the deviation metric—Gextra bounds the cumulative normalized deviation accepted through the tolerance channels (Equation (11))—and does not depend on anything below. What follows supplies the unit conversion behind the interpreted detection rate of Equation (6): we measure, on sampled steps, how much steering progress one normalized unit of tolerated deviation buys relative to one full-power steering update. The comparison is a same-state one: both replacements are evaluated from the same pre-step optimizer state, the tolerated set being a ball centered at the honest aggregate and the fullpower set the clipping ball. It is not a post-center construction in which a full-power update is applied after the trusted step. For a tested step, let gh denote the honest clipped aggregate the verifier recomputes for that step. We compare two replacements of gh inside the same training step, differing only in the set the forged gradient is drawn from: tolerated:

∥g − gh ∥2 ≤ rC,

full-power:

Transfer to the deployment trajectory family. The conversion above is measured on the attack-evaluation trajectories, whose setup differs from the deployment trajectory family (Section 6). We use b κeq only as an empirical scale when interpreting the deployed Gextra budget and do not claim it invariant across trajectory families; establishing that would require repeating this calibration on the deployment trajectories.

∥g∥2 ≤ C.

The tolerated feasible set used here omits the always-on aggregate-norm constraint ∥g∥2 ≤ C(1 + 10−4 ) of the deployed protocol and is therefore a relaxation of the deployed attacker’s feasible set; the comparison is consequently conservative in favour of the attacker. At the tested states, even after adding the largest calibrated tolerance radius, the resulting norm remains far below the deployed cap, so the omitted constraint is inactive on these points; no general claim beyond the tested states is made. Both are solved by the same steering optimizer from the same parameters and optimizer state, and both are scored against the same realized DP noise, so the comparison is a matched pair rather than two separate executions; neither is an additional deployed step. Writing Pttol (r) and Ptfull for the progress each makes on the steering distance relative to the honest step, define Rt (r) =

Pttol (r) , rPtfull

E.4

Value-Aware Checking

The uniform protocol uses pt = p for every iteration. More generally, let pt ∈ [0, 1] be a hidden checking probability selected before training from a public schedule satisfying the T expected-work constraint ∑t=1 pt = pT . For a fixed set F of full malicious deviations, Pdet (F ) = 1 − ∏ (1 − pt ), t∈F

so if certain training phases are known to provide larger attack value, the verifier can assign larger pt to those phases while preserving the same expected number of checks, and an adversary that concentrates its full deviations on those phases faces a correspondingly higher detection probability. This extension concerns the full-deviation detection channel only. All numerical budgets Gsub , Gamb , and Gextra reported in Table 7, and all experiments in this paper, use the uniform setting pt = p; extending the numerical-tolerance accounting to nonuniform pt is outside the present evaluation.

b κeq = max Rt (r). (t,r)∈Ttest

Over nine RoBERTa checkpoints spanning t/T = 0.1 to 0.9, three GPT-2 states on E2E at t/T = 0.20, 0.50 and 0.90, and nine radii from 10−5C to C, every measured Rt (r) exceeds one by less than 10−3 , which is why the maximum rather than the mean is the quantity reported: it is the direction that favours the adversary. We obtain b κRoBERTa = 1.00064 eq LM b and κeq = 1.00038. Thus, over the tested states and radii, a tolerated normalized radius r has approximately r times the steering effect of the corresponding same-state full-power comparator; this same-state ratio is the conversion used in Sections 6.4 and 7.1. The measurement covers the sampled steps, radii, and model families of Ttest , and the conversion in Equation (6) further treats tolerated and full-power deviations as additive and reads the same-state ratio as a step count;

E.5

Status of the Guarantee

(Gextra , βextra ) is an analytic high-probability security guarantee on the cumulative normalized deviation routed through the numerical-tolerance channels, under the stated protocol assumptions; Gextra alone is not a deterministic cap. The guarantee is stated relative to the adaptive trusted reference selected by the verifier on each step: it does not require the TEE FP32 and FP64 executions to define one canonical numerical trajectory, the difference between these two trusted references is not itself charged as adversarial deviation, and the bound concerns the GPU’s deviation from the reference the protocol 28

Table 8: Status of the numerical-tolerance quantities. Quantity

Status

Main dependency

1 − (1 − p)Mfull

analytic

(Gextra , βextra )

analytic highprobability bound trajectoryconditional estimate model-based calibration check

Mfull full deviations under independent hidden Bernoulli checking Bernoulli protocol, frozen parameters, hidden coins exact trajectory, Bernoulli law; the only false-abort estimate representativeness assumption; same trajectories as the calibration one held-out trajectory (qqp-large replay); otherwise untested same-state steering-scale calibration, additive stepequivalent reading; measured on attack-experiment trajectories, transfer to the deployed pair assumed interpreted full-step detection term; not a formal theorem

qFA on the held-out census trajectory in-sample check from systematic samples future-run qFA

empirical generalization

Gextra → full-power steps

empirical interpretation

Equation (6)

empirical interpretation

which allows the GPU to reconstruct the same DP noise and maintain a synchronized local model and optimizer state. Since the seed size is negligible compared with SG , the steady-state communication cost per iteration can be approximated as iter Tcomm ≈

where Band is the effective bandwidth of the GPU–TEE communication path and Tsync captures fixed synchronization costs such as message notification, TEE-boundary crossing, and buffer management. Since the transmitted gradient is an aggregate over the trainable parameters, the communication cost scales with the number of trainable parameters rather than with the batch size or the number of per-example gradients. The GPU-side training, GPU–TEE communication, and TEE-side DP operations form the latency-critical execution path. Let iter iter iter Titer = Ttrain + Tcomm + TDP

denote the average wall-clock time of this path for one iteration. For a run of T iterations, the main training pipeline therefore requires approximately NTiter .

uses to adjudicate that submission. The chain is calibration data → frozen verifier parameters → (Gextra , βextra ): the first arrow is empirical parameter selection, the second is analytic security accounting, so a poor estimate of the honest discrepancy distribution may make the chosen parameters unsuitable for honest availability but does not invalidate the bound evaluated at the parameters actually frozen. Table 8 summarizes the status of each quantity used in this appendix.

F

SG + Tsync , Band

Hungry updating with deferred verification. Under hungry updating, verification is removed from the latency-critical training path. Suppose each iteration is independently selected for checking with probability p, and let Titer-veri denote the average work required to complete the full verification procedure for one checked iteration. Assume that the TEE runs n verification workers in parallel, in addition to the thread serving the main training path. On average, one verification task is generated every 1/p training iterations. Hence, the verification workers receive work at an average rate corresponding to

Efficiency Analysis

The running time of our protocol consists of four main components: (1) GPU-side training time, including forward and backward propagation and clipping7 ; (2) communication time between the GPU and the TEE; (3) TEE-side DP operations, such as noise generation and model update; and (4) TEE-side verification work incurred by probabilistic checking. We use Ttrain , Tcomm , TDP , and Tveri to denote the corresponding total costs.

pTiter-veri units of verification work per training iteration. With n workers, the verification system can process one check every Titer-veri /n units of wall-clock time. Therefore, the verification pipeline can keep pace with training when pTiter-veri ≤ Titer , n

Communication and the main training path. Under our communication-efficient split execution, the GPU sends only the clipped-and-averaged gradient to the TEE, rather than all per-example gradients. Let SG denote the size of this transmitted gradient. After the GPU-submitted gradient is committed, the TEE sends only a short random seed back to the GPU,

(12)

or equivalently, n≥

pTiter-veri . Titer

When Equation (12) holds, verification work can be largely hidden behind the main training pipeline. Training may temporarily run ahead of verification, but a run is considered

7While gradient clipping is a key component of DP-SGD, it is also commonly used in non-DP training to mitigate exploding gradients and improve training stability.

29

complete only after all verification tasks generated during training have finished successfully. Hence, Ttotal = T Titer + Tdrain , where Tdrain denotes the time required to drain any remaining verification tasks after the final training iteration. In the stable regime, Tdrain is small, and therefore Ttotal ≈ T Titer . On the other hand, if pTiter-veri > Titer , n verification tasks are generated faster than the workers can process them, and a backlog accumulates. Under a steadystate approximation, the end-to-end running time becomes   pT Titer-veri Ttotal ≈ max T Titer , . n In the verification-bottleneck regime, this reduces to Ttotal ≈

pT Titer-veri . n

Thus, hungry updating converts verification from a synchronous per-check latency into a background throughput requirement. When sufficient verification parallelism is available, most verification work overlaps with the main training pipeline, resulting in only a small end-to-end overhead.

G

MIA parameters

The parameters of the MIA experiments are shown in Table 9.

Table 9: Hyperparameters of the membership-inference experiments. IMIA

SHAPOOL

Hyperparameter

Setting Hyperparameter

Setting

Number of imitative models Imitative-out total epochs Warm-up epochs Imitation epochs Pivot fine-tuning epochs Shadow / imitation batch size Shadow / imitation learning rate Shadow / imitation dropout Pivot samples per class

10 100 80 20 20 256 0.1 0 100

10 100 128 0.1 3 64 0.1 5 0.5 0

Number of shadow models Shadow pre-training epochs Shadow pre-training batch size Shadow pre-training learning rate Fine-tuning epochs Fine-tuning batch size Fine-tuning learning rate Number of experts MoE ratio Dropout

30

Record · ID 978350 · SHA-256 08e922a3a2d2ac4a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.