RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning Kwunhang Wong1,2,∗ , Jichang Yang1,2,∗ , Karl M.H. Lai2 , Hegan Chen1,2 , Songqi Wang1,2 , Wei Xuan1 , Ning Lin2,4,† , Han Wang2 , Xiaojuan Qi2,† and Zhongrui Wang1,3,†
arXiv:2607.18169v1 [cs.CR] 20 Jul 2026
Abstract Edge Artificial Intelligence of Things (AIoT) systems often collect sensitive data in situ, raising serious privacy concerns. Resistiveswitching random-access memory (RRAM) is an attractive substrate for efficient AIoT thanks to its multi-bit storage and compute-inmemory (CiM) capabilities, while its inherently stochastic write behavior provides a natural source of randomness that can be leveraged for differential privacy (DP) protection. Yet how to transform this device-level randomness– typically viewed as detrimental to accuracy– into a principled randomized mechanism while preserving model utility remains underexplored. We propose RRAM-DP, a hardware–algorithm co-design that relaxes RRAM write–verify operations to inject calibrated noise for inherently (𝜀, 𝛿)-DP with formal DP analysis; together with pretraining techniques, it renders a novel private, high-utility CiM-training paradigm. On CIFAR-10/100, STS-B, and SST-2, RRAM-DP-SGD incurs at best only a 3.8% accuracy drop at (𝜀=2,𝛿=O(1/n))-DP relative to non-private SGD. At the same privacy level, RRAM-DP-SGD delivers up to 57×& 3.2× energy savings and 2.7×& 1.8× speedups over A100 and DiVa-GEMM, respectively. These results point towards efficient, privacy-preserving in-memory training on RRAM at the edge.
CCS Concepts • Hardware → Memory and dense storage; • Security and privacy → Formal methods and theory of security.
Keywords RRAM, Memristor, AI accelerator, Compute-in-Memory, Differential Privacy, Edge AIoT, Device Variability
1
Introduction
Edge Artificial Intelligence of Things (AIoT) systems increasingly operate in situ sensing and learning from local, privacy-sensitive environments to personalize services across modalities and tasks [45]. This shift heightens well-documented risks: seemingly anonymized records can be re-identified [31], and modern deep neural networks (DNNs) memorize and leak training data via membership inference attacks (MIAs) [7] and training data extraction attacks [6, 8]. Differential privacy (DP) provides a principled defense by adding calibrated noise to query outputs, masking any single individual’s contribution and thereby offering rigorous protection against databaselevel attacks [15]. DP-Stochastic Gradient Descent (SGD) extends DP guarantees to integrating noise injection in SGD optimizer, enabling privately trained DNNs with (𝜀, 𝛿)-DP [1]. However, despite their widespread use, existing DP and DP-SGD frameworks still face critical challenges in AIoT applications as detailed below.
Probability Density
1 ACCESS – AI Chip Center for Emerging Smart Systems, InnoHK Centers, Hong Kong Science Park, Hong Kong; 2 Department of Electrical and Computer Engineering, The University of Hong Kong, Hong Kong; 3 School of Microelectronics, Southern University of Science and Technology, Shenzhen, China; 4 Department of Electrical Engineering, City University of Hong Kong, Hong Kong ∗ Equal contribution to this work; † Corresponding authors: [email protected]; [email protected]; [email protected]
0.8
HRS
0.6
Gradient
0.01
0.4
RESET
0.2 0.0
0.03
SET 0
write
LRS
20 40 60 80 Conductance (μS)
100
Privacy attacks: MIAs, Training Data Extraction, etc.
Data statistics ratio≤eε
M(D)
M(D’) ζ(.)
Crossbars (ε,δ)-DP Guarantee RRAM-DP AI Accelerator Analog Storage Compute-in-Memory
Figure 1: (LEFT) Histogram of cycle-to-cycle variation for an RRAM device between high resistance state (HRS) and low resistance state (LRS) with 300 iterative SET-RESET at fixed operation voltages. (RIGHT) Such noise can provide (𝜀, 𝛿)-DP for storing sensitive data statistics/model parameters. To date, most DP mechanisms are realized in digital hardware– namely CPUs/GPUs/TPUs and dedicated digital accelerators [32, 42]– which raises several tensions intersected at efficiency and privacy for edge AIoT: (i) pseudo-random noise sampling in digital platform has finite precision with rounding errors, which is susceptible to floating-point error attacks [21, 30]; (ii) repeated data movement between processors and memory for real-time randomized mechanism are energy-expensive in local DP [37] (iii) existing DP-SGD implementation libraries suffer 2 × −1000× in time and space complexity compared to normal SGD training [4], which are especially problematic in resource constrained edge. Resistive-switching random-access memory (RRAM) is promising solution to address these challenges: First, RRAM’s inherently stochastic programming offers a native source of entropy for true random number generation. Second, write operations of RRAM, originally regarded as a reliability problem, are naturally an in-place randomization with analog storage capability [39]. Thirdly, the ohmic nature of RRAM naturally allows efficient multiply-accumulates (MACs) computation, where RRAM-based Compute-in-memory (CiM) accelerator provides powerful in-memory vector-matrix multiplication. Prior DP work explored memristor noise largely at a theoretical level, under the assumption of idealized i.i.d. Gaussian perturbations [18]. In practice, RRAM noise is non-identical, and practically difficult to be controlled across devices and cycles for different privacy levels–whether how device-realistic randomness can be calibrated to deliver formal DP guarantees [15] remains open. In addition, CiM-training usually underperforms digital accelerators due to noise accumulating in write operations [20]. A scalable CiM-training paradigm is another challenge. To this end, we introduce RRAM-DP, which is a co-design that quantifies device randomness into privacy levels, along with pretraining techniques to boost CiM-training utility. To our knowledge,
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
K. Wong et al.
we propose the first RRAM-based randomized mechanism with (𝜀, 𝛿)DP for both on-device data storage and SGD executed directly in RRAM through formal 𝜇-Gaussian DP (𝜇-GDP) analysis grounded in device statistics. For hardware, we realize a robust noise source at program time by relaxing write-verify cycles with only a single control variable 𝜏. The overview of RRAM-DP is shown in Fig. 1 and we summarize our contributions as follows: • First RRAM-based randomized mechanism. Contrary to digital hardware with finite randomness and efficiency issues, we calibrate the privacy level of RRAM write operations for in situ true noise addition (𝜀, 𝛿)-DP mechanisms, namely RRAM-DP mechanism and RRAM-DP-SGD. • A robust DP noise sampling from RRAM. Challenges such as non-linear filament growth and fabrication non-idealities across devices are addressed by write-verify sampling and Lindeberg-Central Limit Theorem (CLT) design, calibrating steady noise with only one parameter. • Noise-resilient CiM-training paradigm. To mitigate the reliability degradation of noisy write operation, we connect pretraining techniques inspired by DP-SGD to in situ RRAMbased CiM-training. Results show RRAM-DP-SGD incurs ≤∼7% accuracy drop at 𝜀=5 compared to ideal training. • State-of-the-art efficiency of DP-SGD accelerator. Energy expensive DP operations can be locally processed and stored by RRAM-DP-SGD accelerator to prevent privacy leakage. Simulations show a 54×-57× and 3×-3.2× energy savings; 2.5×-2.7× and 1.7×-1.8× speedup compared to A100 and digital DP-SGD accelerator DiVa-GEMM [32].
2
2.2
Standard RRAM cell
Definition 2.1 ((𝜀, 𝛿)-Differential Privacy[15]). Given a randomized mechanism M : D → O is (𝜀, 𝛿)-DP if for any neighbouring dataset pairs 𝐷, 𝐷 ′ ∈ D from the input domain, and for any measurable output subset 𝑆 ⊆ O: Pr[M (𝐷) ∈ 𝑆] ≤ 𝑒 𝜀 Pr[M (𝐷 ′ ) ∈ 𝑆] + 𝛿 If 𝛿 = 0, M satisfies 𝜀-DP. The parameter 𝜀 measures the maximum privacy loss of M. While 𝛿 allows for a small probability of 𝜀 failure, 𝛿 should be smaller than 1/|𝐷 | to prevent privacy breach for tail examples [1, 15]. This DP guarantee cannot be weakened by analyzing the DP mechanism output unless there exists additional access to the original data. Definition 2.2 (Δ(𝜁 ) Sensitivity[14]). Let 𝜁 : D → O𝑑 be any 𝑑-dimensional query function. The 𝐿2 sensitivity of function 𝜁 for any neighbouring 𝐷, 𝐷 ′ ∈ D can be defined: Δ(𝜁 ) ≜ max′ ∥𝜁 (𝐷) − 𝜁 (𝐷 ′ )∥ 2 𝐷,𝐷
2.3
Gaussian Differential Privacy
GDP has the tightest known bounds on composition for mechanisms that closely align with Gaussian noise, such as DP-SGD [3]. Since hardware RRAM noise modeling is inherently noisy and dynamic compared to software DP deployment, we choose 𝜇-GDP as an analytical tool for tightest privacy estimation of calibrated RRAM noise. By definition, 𝜇-GDP implies a lossless conversion towards a collection of (𝜀, 𝛿)-DP guarantees. Definition 2.3 (Trade-off Function 𝑇 [13]). Let 𝜙 be a measurable rejection rule for distinguishing distributions P and Q on the same space; 𝛼𝜙 and 𝛽𝜙 be the probabilities of an adversary making Type I and Type II errors, respectively. The trade-off function 𝑇 P | | Q (𝛼) : [0, 1] → [0, 1] is defined as: 𝑇 P | | Q (𝛼) ≜ 𝑖𝑛𝑓 {𝛽𝜙 : 𝛼𝜙 ≤ 𝛼 }
RRAM uses analog conductance 𝐺 to denote a real number 𝜃 via a mapping function 𝑔 : [𝜃 min, 𝜃 max ] →[𝐺 min, 𝐺 max ]. An example mapping function could be: 𝑔(𝜃 ) = 𝐺 min + (𝜃 − 𝜃 min ) ·
Differential Privacy
DP is a robust privacy guarantee (usually 𝜀 ≤10 [19]) to protect all records in a database by estimating the worst-case output change from the inclusion or exclusion of one record in the database. The technical formulation of DP is as follows.
Preliminaries
Problem Statement. When AIoT parameters are not duly handled by privacy-preserving algorithms, an adversary can infer useful information through various privacy attacks. The least revealing attack is MIAs, which formulate a binary hypothesis test to decide if a datapoint exists in the training database [2, 7]. Model inversion attacks infer features that characterize each output class [17], making it possible to reconstruct training data, for example, if all class members depict the same person in a facial recognition model. The worst revealing attack is data extraction attack; generative models such as large language models (LLMs) [6] and diffusion [8] can memorize training data from which adversarially crafted prompts can expose individual’s bank account number, images, etc. These privacy attacks call for privacy-preserving techniques, namely randomized DP mechanisms, to protect privacy-sensitive edge parameters.
2.1
Motivations. We address the challenges of non-ideal noisy RRAM write by experimentally measured variability, and propose a principled method for sampling and calibrating device-realistic noise to satisfy the formal definition of DP in this work. By linking device-level statistics to privacy parameters, our approach converts unavoidable hardware non-idealities into controlled privacy guarantees without architectural modification.
𝐺𝑟𝑎𝑛𝑔𝑒 𝜃 𝑟𝑎𝑛𝑔𝑒
(1)
where 𝐺𝑟𝑎𝑛𝑔𝑒 = 𝐺 max − 𝐺 min , 𝜃 𝑟𝑎𝑛𝑔𝑒 = 𝜃 max − 𝜃 min . However, writing a target conductance to RRAM suffers from inevitable randomness due to the underlying electrochemical ion migration [44]. More specifically, there exist cycle-to-cycle variation and device-to-device variation that constitute a complex and dynamic RRAM noise model on a crossbar discussed in sec.3.
The infimum is taken over all measurable rejection rules 𝜙. The function 𝑇 P | | Q characterizes the optimal trade-off between Type I and Type II errors for distinguishing P from Q. Definition 2.4 (𝜇-Gaussian Differential Privacy [13]). Given a randomized mechanism M is 𝜇-GDP (also known as 𝑓 -DP, where 𝑇 (M (𝐷), M (𝐷 ′ )) ≥ 𝑓 , 𝑓 is gaussian-parameterized) if for any neighbouring dataset pairs 𝐷, 𝐷 ′ ∈ D: 𝑇 (M (𝐷), M (𝐷 ′ )) ≥ 𝑇 (N (0, 1), N (𝜇, 1)) = 𝐺 𝜇 The exact solution of 𝐺 𝜇 (𝛼) = Φ(Φ−1 (1 − 𝛼) − 𝜇) is derived by the Neyman–Pearson lemma [16], which states the most powerful test
RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
Theorem 2.1 ((𝜀, 𝛿)-DP Conversion[13]). For 𝜀 ≥ 0, a randomized mechanism M is 𝜇-GDP if and only if M is (𝜀, 𝛿)-DP that follows: 𝜀 𝜇 𝜀 𝜇 𝛿 (𝜀) ≜ Φ(− + ) − 𝑒 𝜀 Φ(− − ) 𝜇 2 𝜇 2 We denote satisfying 𝜇-GDP is the same as satisfying (𝜀, 𝛿)-DP in the following paragraphs for simplicity.
3
RRAM-DP Formulation
Protection Overview. RRAM-DP is a co-design framework formulated under an RRAM-based randomized mechanism at one output parameter level to satisfy 𝜇-GDP. Firstly, a 𝑘-device configuration of RRAM crossbars is designed in which every analog output parameter is programmed on 𝑘 devices with shared write control and read-averaging circuitry. Secondly, the cycle-to-cycle variation is calibrated by the convergence of a write-verify algorithm at the target conductance within an adjustable absolute error bound 𝜏. For statistics storage, write operations inject noise with RRAM-DP mechanism to provide DP guarantee. For DNNs, write operations in RRAM-DP-SGD provides DP guarantee for trained models.
3.1
Additive Noise Strategy
To calibrate noise, this work follows a classic write-verify method that writes RRAM iteratively with a feedback loop control [29]. As described in Fig. 2, this method reads an RRAM conductance in the current cycle and proceeds to the next SET/RESET cycle until hitting within stopping threshold. This method can steadily capture variation due to a bounded sampling region, disregarding overshooting/undershooting problems from the non-linear filamentary growth [20]. The SET/RESET program voltage adjustment rules vary based on device characteristics heuristically. Cycle-to-cycle Variation. Assume an absolute conductance error Δ𝐺 = 𝐺 𝑤 − 𝐺𝑡 is collected from differencing written conductance 𝐺 𝑤 and target conductance 𝐺𝑡 . We first denote that there are 𝑁 devices for storing 𝑁 parameters. Then, the final cycle-to-cycle variation of Δ𝐺𝑖 for each device 𝑖 sampled at the converged writeverify algorithm follows an observed distribution Dist(𝜏, 𝐺𝑡 ), which is dependent on the algorithmic stopping threshold 𝜏 and target conductance 𝐺𝑡 = 𝑔(𝜃 𝑖 ): 𝐺𝑟𝑎𝑛𝑔𝑒 Δ𝐺𝑖 = Δ𝜃 𝑖 · , 𝑖 = 1, ..., 𝑁 (2) 𝜃 𝑟𝑎𝑛𝑔𝑒 𝑔𝑖𝑣𝑒𝑛 Δ𝐺𝑖 ∼ Dist(𝜏, 𝐺𝑡 ), ∀𝜏 ∈ R+, 𝐺𝑡 ∈ [𝐺 min, 𝐺 max ] Eq. 2 maps the program error of an RRAM device that contributes to numerical error by 𝐺𝑟𝑎𝑛𝑔𝑒 : 𝜃 𝑟𝑎𝑛𝑔𝑒 projection rule stated in Eq. 1. As the program voltage aims exactly towards 𝐺𝑡 , this usually gives a favourable DP noise source for E[Δ𝐺𝑖 ] ≈ 0. This work will prove the relaxation of the 𝐺𝑡 condition by our experiments afterwards for a simple yet effective noise calibration. Then, we can express the precise variance control only by 𝜏 for (𝜀, 𝛿)-DP requirement. Denoting the parameter error Δ𝜃 as a random variable in Eq. 3, 𝜎𝑖 (Δ𝜃 ) = 𝜎𝑖 (Δ𝐺) ·
𝜃 𝑟𝑎𝑛𝑔𝑒 , 𝑖 = 1, ..., 𝑁 𝐺𝑟𝑎𝑛𝑔𝑒
G max
LRS
Readout Conductance Writing Ends !!
G target Start @HRS
RES #2 SET #1
RO #1
Stopping Threshold
Converge
G min
Cycle SET RESET Readout
SET/RESET Program SL BL Voltage Voltage
of differentiating two simple hypotheses is a likelihood ratio test. Φ denotes the cumulative density function of the Gaussian distribution, which applies to sec. 3.3.
RES #3 RO #2
RO #3
SET #4
RO #4
SET #5
RO #5
End if stopping precision met
Figure 2: Control flow diagram for a classic RRAM write-verify algorithm. Voltage pulses into bit lines (BLs)/source lines (SLs) control SET/RESET, respectively. the hyperparameter 𝜃 𝑟𝑎𝑛𝑔𝑒 increases or 𝐺𝑟𝑎𝑛𝑔𝑒 decreases. As Fig. 2 shows, more cycles are required for the convergence towards a smaller stopping threshold with smaller variance. This suggests the challenge of device-calibrated RRAM-DP lies in noise minimization rather than maximization. Device-to-device Variation. The non-ideal fabrication process, such as etching-induced damage and thickness of switching film [20], causes each RRAM device to have a non-identical stochastic distribution of Δ𝐺𝑖 ∼Dist(𝜏, 𝐺𝑡 ). Ideally, the cycle-to-cycle variation for every individual device should be independently modelled to ensure necessary noise is added without violating a strict (𝜀, 𝛿)-DP guarantee. Since auditing all devices is non-trivial, notably in very large crossbars, this work develops RRAM-DP mechanism for device-todevice noise modeling relaxation below.
3.2
Non-ideal Noise with Multi-device Representations
3.2.1 Smoothen Device-to-device Spread. We denote 𝑘 × 𝑁 devices e𝑖 ∼Dist(𝜏, 𝐺𝑡 ) to address devicefor 𝑁 parameters analog storage as Δ𝐺 to-device variation. Increasing 𝑘 reduces spread and smoothens outlier RRAM devices for a steady, consistent noise source at the cost of hardware. 3.2.2 Unify Device-to-device Gaussian Shape. A 𝑘-device configuration design with averaged sum output satisfies the Lindeberg-CLT as 𝑘→∞ [23], as summing any independent random variable with bounded distribution and variance eventually converges to Gaussian. We formulate RRAM-DP randomized mechanism for the statistical behaviour of a global non-uniform Gaussian write below: Theorem 3.1. For some dataset 𝐷 ∈ D, we define an arbitrary query function 𝜁 : D → O𝑑 and a sufficiently large k-device representation to write one parameter 𝜃 𝑖 ∈ 𝜁 (𝐷). Then, there exists an in situ RRAM-DP mechanism R on the averaged sum output that satisfies 𝜇𝑟 -GDP where: R (𝐷) ≜ 𝜁 (𝐷) + N (0, diag(𝜎𝑖 (Δ𝜃 ) 2 ))
(4)
𝐺 e ≥ Δ(𝜁 ) · range , ∀𝑖 = {1, 2, ..., 𝑁 } such that 𝜎𝑖 (Δ𝐺) 𝜇𝑟 𝜃 range
(3)
The standard deviation function 𝜎𝑖 (.) on Δ𝜃 implies the upper bound of RRAM-controlled parameter error can go unbounded as
Proof. Assume writing 𝜃 𝑖 is strictly independent Gaussian due to CLT and physics randomness, querying 𝜁 for two neighboring datasets 𝐷 ∼ 𝐷 ′ will be two multivariate Gaussian with 𝑃 ∼ N (𝑚, Σ),
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
K. Wong et al.
𝑄 ∼ N (𝑚 ′, Σ) for writing all 𝜃 𝑖 . By Neyman-Pearson lemma of 𝜇GDP in definition 2.4, a likelihood ratio test Λ(𝑥) is considered for the privacy loss random variable: Λ(𝑥) =
exp[− 12 (𝑥 − 𝑚)𝑇 Σ −1 (𝑥 − 𝑚)] 𝑓𝑃 (𝑥) = 𝑓𝑄 (𝑥) exp[− 12 (𝑥 − 𝑚 ′ )𝑇 Σ −1 (𝑥 − 𝑚 ′ )]
1 1 = exp[− (𝑥 − 𝑚)𝑇 Σ −1 (𝑥 − 𝑚) + (𝑥 − 𝑚 ′ )𝑇 Σ −1 (𝑥 − 𝑚 ′ )] 2 2 ′ Δ 𝑚 +𝑚 )] = exp[Δ𝑇 Σ −1 · ( + 𝑧)] = exp[Δ𝑇 Σ −1 · (𝑥 − 2 2 where Δ = 𝑚 − 𝑚 ′ ; 𝑥 ∼ 𝑃, 𝑧 ∼ N (0, Σ). Denote Δ𝑇 Σ −1𝑧 ∼ N (0, Δ𝑇 Σ −1 Δ). The privacy loss of RRAM-DP mechanism R can be shown as log Λ(𝑥)|𝑃 ∼ N ( 12 Δ𝑇 Σ −1 Δ, Δ𝑇 Σ −1 Δ). Similarly, 𝜇-GDP has a privacy loss of log Λ(𝑥)|𝑥∼N (𝜇,1) ∼ N ( 12 𝜇 2, 𝜇 2 ). √ Therefore, the goal is to identify the maximum 𝜇 = Δ𝑇 Σ −1 Δ ∀𝐷 ∼ 𝐷 ′ such that ||Δ|| 2 ≤ Δ(𝜁 ) to declare 𝜇-GDP for the process. This yields a Rayleigh-quotient maximization problem:
Theorem 3.2. Let C𝑝 (R) ⊗𝑇 denote the 𝑇 -fold composition of RRAM-DP mechanism R under uniform subsampling rate 𝑝 = 𝐵𝑛 . ′ ′ Define 𝜎min ≤ inf 1≤𝑡 ≤𝑇 𝜎min,𝑡 for all iteration 𝑡. Then, in the asymp√ totic regime 𝑝 𝑇 → 𝑐 as 𝑇 → ∞, the RRAM-DP-SGD algorithm of ′ ) is 𝜇 -GDP: C𝑝 (R) ⊗𝑇 with per-step noise multiplier diag(𝜎𝑖,𝑡 𝑟 √︃ ′ −2 ′ −1 ′ −1 ) + 3Φ(−0.5𝜎min ) − 2) (5) 𝜇𝑟 = 𝑝 · 2𝑇 · (𝑒 𝜎min · Φ(1.5𝜎min 2𝐶𝜂 𝐺𝑟𝑎𝑛𝑔𝑒 ′ e ≥ 𝜎min · , ∀𝑖 = {1, 2, . . . , 𝑁 } such that 𝜎𝑖 (Δ𝐺) · 𝐵 𝜃 𝑟𝑎𝑛𝑔𝑒 Proof. In iteration 𝑡, given lines 6 and 8 for any two neighbouring datasets 𝐷 ∼ 𝐷 ′ , the per-iteration gradient update 𝜁 (𝐷) = 𝜂 Í 2𝐶𝜂 𝑗 ∈𝐼𝑡 𝑔¯𝑥 𝑗 has sensitivity bound Δ(𝜁 ) = 𝐵 . In line 9, RRAM-DP𝐵 ′ )2 , SGD adds calibrated noise Δ𝜃 ∼ N (0, Σ𝑡 ) , Σ𝑡 = diag (Δ(𝜁 ) 𝜎𝑖,𝑡 ′ so that the 𝑖-th coordinate noise standard deviation is Δ(𝜁 ) 𝜎𝑖,𝑡 . Let ′ ′ . By theorem 3.1 (RRAM-DP mechanism GDP guar𝜎min,𝑡 = min𝑖 𝜎𝑖,𝑡 antee), the mechanism at iteration 𝑡 is 𝜇𝑡 -GDP with 𝜇𝑡 = 𝜎 ′ 1 . Since min,𝑡
max | |Δ| | 2 ≤Δ(𝜁 )
Δ Σ Δ = Δ(𝜁 ) 𝜆max (Σ ) 𝑇
−1
2
−1
where 𝜆max (Σ −1 ) is√︁the largest eigenvalue of Σ −1 . This gives Δ(𝜁 ) 𝜆max (Σ −1 ) ≤ 𝜇𝑟 to claim 𝜇𝑟 -GDP. With Σ = 2 . Therefore, we diag(𝜎𝑖 (Δ𝜃 ) 2 ), 𝜆max (Σ −1 ) = max𝑖 (1/𝜎𝑖2 ) = 1/𝜎min 𝐺 Δ(𝜁 ) Δ(𝜁 ) ˜ ≥ have 𝜎min (Δ𝜃 ) ≥ ⇒ 𝜎min (Δ𝐺) · 𝑟𝑎𝑛𝑔𝑒 by substituting 𝜇𝑟
𝜇𝑟
min,𝑡
□
RRAM-DP-Stochastic Gradient Descent
Most RRAM CiM-accelerators only function as inference instead of training due to accumulative write variance during weight update [20, 40]. Such RRAM write noise has a natural fit with the iterative noise addition in DP-SGD, which protects training data from being released in the privately trained model with (𝜀, 𝛿)-DP at the edge. We formulate RRAM-DP-SGD over the composition of 𝑇 iterative RRAM-DP mechanism on the gradient function with private subsampling amplification in Algo. 1: Algorithm 1 RRAM-DP-SGD 1: Input: Private examples 𝐷 = {𝑥 𝑗 | 𝑗 = 1, 2, . . . , 𝑛},
Loss function L (𝜃, 𝑥). Parameters: learning rate 𝜂, noise model ′ , clip norm 𝐶, batch size 𝐵, iterations 𝑇 𝜎min 2: for 𝑡 = 0, 1, . . . ,𝑇 − 1 do 3: Random-subsample 𝐷𝑏 uniformly where 𝐷𝑏 ⊂ 𝐷, 𝐷𝑏 = {𝑥 𝑗 | 𝑗 ∈ 𝐼𝑡 }, |𝐼𝑡 | = 𝐵 4: for each 𝑗 ∈ 𝐼𝑡 do 5: Compute gradient: (RRAM) 𝑔𝑥 𝑗 ← ∇𝜃 L (𝜃 𝑡 , 𝑥 𝑗 ) 6: Clip gradient: 𝑔¯𝑥 𝑗 ← 𝑔𝑥 𝑗 · min(1, 𝐶/∥𝑔𝑥 𝑗 ∥ 2 ) 7: end for 8: Average gradient and descend: Í 𝜃¯𝑡 +1 ← 𝜃 𝑡 − 𝜂 · 𝐵1 𝑗 ∈𝐼𝑡 𝑔¯𝑥 𝑗 9: Projection (|𝜃¯t+1 | = 𝑁 ) & Write: (RRAM) e ≥ 𝜎 ′ · 2𝐶𝜂 · 𝐺𝑟𝑎𝑛𝑔𝑒 , ∀𝑖 = {1, 2, , ..., 𝑁 } 𝜎𝑖 (Δ𝐺) min 𝐵 𝜃𝑟𝑎𝑛𝑔𝑒 10: Estimate privacy 𝜇𝑟 -GDP at iteration 𝑡 + 1 (Eq. 5) 11: end for 12: Output: 𝜃𝑇 with privacy budget (𝜀, 𝛿)-DP conversion
min
define the actual worst-case per-step GDP parameter 𝜇max : 𝜇 max ≜ sup 𝜇𝑡 ≤ 1≤𝑡 ≤𝑇
𝜃𝑟𝑎𝑛𝑔𝑒
𝜆max (Σ −1 ), which implies the conductance inequality ∀𝑖.
3.3
sampling all device true distributions is non-trivial, we model an ′ empirical statistical lower bound function with 𝜎min (see sec. 4.3). 1 ′ ′ Then for all 𝑡, we have 𝜎min,𝑡 ≥ 𝜎min ⇒ 𝜇𝑡 = 𝜎 ′ ≤ 𝜎 ′1 . We 1 ′ 𝜎min
Let 𝐺 𝜇 denote the trade-off function of a 𝜇-GDP Gaussian mechanism. The composition rule for GDP [13] states that composing 𝜇𝑡 -GDP mechanisms gives 𝐺 𝜇1 ⊗ · · · ⊗ 𝐺 𝜇𝑇 = 𝐺 √︃ 2 2 . 𝜇1 +···+𝜇𝑇
Since 𝜇𝑡 ≤ 𝜇max for all 𝑡, v u t𝑇 √ √︃ ∑︁ √ 𝑇 2 2 𝜇𝑡 ≤ 𝑇 𝜇max = 𝑇 𝜇max ≤ ′ 𝜎 min 𝑡 =1 Therefore, the 𝑇 -fold (non-subsampled) composition satisfies 𝐺 𝜇1 ⊗ · · · ⊗ 𝐺 𝜇𝑇 = 𝐺 √︃Í𝑇
2 𝑡 =1 𝜇𝑡
⪰ 𝐺 √𝑇 𝜇max ⪰ 𝐺 √𝑇 /𝜎 ′
min
Equivalently, we can model each iteration as a 𝜇 ′ -GDP mechanism with 𝜇 ′ = 𝜎 ′1 as the base mechanism with trade-off function min
𝐺 𝜇 ′ ; thus, the 𝑇 -fold composition C𝑝 (R) ⊗𝑇 with privacy amplification convergence through batch subsampling rate 𝑝 in the as√︃ √ 2 ymptotic regime is no worse than 𝜇𝑟 = 𝑝 2𝑇 𝜒+ (𝐺 𝜇 ′ ) as 𝑝 𝑇 → constant 𝑐, where 𝜒+2 (𝐺 𝜇 ) = 𝑒 𝜇 Φ(3𝜇/2) + 3Φ(−𝜇/2) − 2 under GDP Lemma 5.3 [13]. Therefore, substituting 𝜇 ′ = 𝜎 ′1 yields 𝜇𝑟 = min √︃ ′ −2 ′ −1 ′ −1 𝑝 · 2𝑇 · (𝑒 𝜎min · Φ(1.5𝜎min ) + 3Φ(−0.5𝜎min ) − 2). □ 2
In Algo. 1, we accelerate RRAM-DP-SGD in an analog-digital computing fashion, where gradient computation (line 5) and noise generation (line 9) can be directly processed by analog RRAM crossbar. Since RRAM noise occurs at parameter level instead of gradient, we reverse the normal gradient descent compute logic of DP-SGD in lines 8 and 9, which causes the privatised gradient sensitivity Δ(𝜁 ) to scale with learning rate 𝜂 and batch size 𝐵. Lastly, we propose an RRAM-DP-SGD accountant (Eq. 5) in theorem 3.2 to rigorously ′ track device non-idealities with 𝜎min after training iteration 𝑡 for line 9. RRAM-DP-SGD is defined under the replacement definition with 2 times the sensitivity of the add-or-remove definition [22].
RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning
Private Input Examples
Length (n) 𝜹 Pretraining Fine-tuning Batch Size (B) Clip Norm (C) LR (𝜂) [𝐺 min, 𝐺 max ] [𝜃 min, 𝜃 max ]
Datasets
STS-B
SST-2
50k 50k 7k 67k 10−5 10−5 10−4 10−5 ImageNet32 [9] ImageNet32 [9] Huggingface [12] Huggingface [27] Entire Model Entire Model LoRA(r=2) LoRA(r=4) 2048 2 1.5 [20, 80] [-1, 1]
2048 1 1.5 [20, 80] [-2, 2]
256 1.5 0.5 [20, 80] [-1, 1]
2048 2 1.5 [20, 80] [-1, 1]
Hardware Accelerator. Our hardware accelerator shown in Fig. 3 borrows a general RRAM-based CiM architecture [38] to implement RRAM-DP-SGD. The system core consists of a global control unit and 512 tiles; each tile contains 16×16 processing elements (PEs). Each PE contains three 32×32 RRAM arrays, where a weight value
RRAM Read/Write Driver
...
PE
...
...
...
PE
PE
...
PE
Post Processor
Noisy Weights
Buffer
... ...
...
... ...
...
... ...
PE
... ... ... ...
PE
Row Selector
PE
DAC
...
ADC
PE
...
I/O Buffer
Global Memory
PE
Mode Selector
BL
Control Unit
... ... WL
...
SL
Column Selector Mode Selector DAC ADC
Figure 3: Architecture of an RRAM-based CiM accelerator for RRAM-DP-SGD training, including the system-level framework and the circuit-level configuration of a single PE for 𝑘 RRAM macros with peripheral circuits. Device #1 Device #2 Device #3
0.4 0.2
Conductance error (μS)
0.0
-2
4 3 2 1 0 -1 -2 -3
0
2/-2 0 2/-2 Conductance Error (μS)
2
0.4
2 Devices 3 Devices Gaussian -3
0 3 Devices
Device #1 Device #2 Device #3
Density
Density
0.6
Table 1: Simulation settings for RRAM-DP-SGD CIFAR100
RRAM-based Store (ε,δ)-DP DNN CiM Core
k - device
4 Experiments 4.1 Experimental Setup Memory Devices. The electrical characterization is performed using a test chip fabricated in a 180 nm process, which integrates a 32×32 one-transistor-one-resistor (1T1R) macro. The memory cells comprise TiN as bottom and top electrodes with a Ta2 O5 /TaO𝑥 dielectric stack. The 1T1R array employs a crossbar topology, with row-shared word lines (WLs) and source lines (SLs) and column-shared bit lines (BLs) for efficient cell addressing. Hybrid Analog-digital System. We perform hybrid analogdigital training from scratch on MNIST (8×8 resized) with a 2-layer MLP network of 760 weights. Weights are stored in analog RRAM arrays, whereas dynamic parameter updates are implemented in a digital computing unit with SRAM buffer. The minimum viable testing uses hyper-parameters 𝛿=10−5 , batch size=2048, clip norm=2, LR=2, [𝐺 min , 𝐺 max ]=[20, 80], [𝜃 min , 𝜃 max ]=[-1, 1]. Simulation Benchmarks. We scale up RRAM-DP-SGD on Opacus [42] with experimentally acquired RRAM noise statistics on the 32×32 1T1R macro for CIFAR-10, CIFAR-100 [24], STS-B, and SST2 [36] datasets. Following [11], we use pretrained models of WRN16-4, WRN-28-10 [43], BERT-Base [12], and RoBERTa-Base [27], respectively. Table 1 shows detailed private training hyper-parameters simulated by NVIDIA RTX 4090 GPU.
CiM Controller
Tile
I/O Interface
The program precision from memristor non-idealities causing significant degradation in network performance (e.g., ≥60% accuracy drop on CIFAR10/100 [26]) is a major bottleneck for memristor-based training [33, 38]. While this work rethinks the formal connection between device-realistic noise and theoretical DP-SGD (noisy training), we find an elegant way of borrowing DP-SGD training techniques to improve utility in CiM-training. Due to the noisy privatised mechanism, DP-SGD has limited utility as compared to normal SDG training. One of the scalable methods for improving DP-SGD utility is through pretraining non-privately on massive public data or synthetic data, and fine-tuning with DP-SGD on precious private data [11, 34, 35]. We believe more scaling techniques from DP-SGD being migrated to improve analog CiM-training are worth future research exploration despite the differences in calibrating devicerealistic noise; and we adopt a pretraining strategy from DP-SGD in this work [11], such as a larger batch size and group normalization layer, to improve general CiM-training utility.
CIFAR10
RRAM Macro
RRAM-DP-SGD Accelerator
Pretraining the CiM-based Training
CPU
3.4
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
-2 -1 0 1 2 3 Theoretical Quantiles
0.3 0.2 0.1 0.0
-2 0 2 Conductance Error (μS)
Figure 4: (TOP) Histogram of 300 samples from 3 independent RRAM devices, sampled at Δ𝐺𝑖=1,2,3 ∼ Dist(𝜏=3µS, 𝐺𝑡 =50µS), respectively. (BOTTOM) Q-Q plot for the averaged sum & Histogram of the 3-device representation. is stored by a 3-device configuration. Along with peripheral circuits such as digital-to-analog converters (DACs), analog-to-digital converters (ADCs), and read/write drivers, it performs energy-efficient in-memory vector-matrix multiplication. Evaluation Metrics. Model utility: We compare test accuracy of RRAM-DP-SGD with non-private SGD training. Minimal performance drop reports good utility. Model training cost: We evaluate FLOPS [41] for one iteration. Hardware performance: We compare energy and time cost over FP32 training on A100 GPU (40GB) at peak 19.5TFLOPS, maximum power 250W; and DiVa-GEMM at peak 29.5TFLOPS, maximum power 21.2W [32].
4.2
Robust Gaussian for 3-Device Representation
We first study the sufficient 𝑘-device representation to generate Gaussian noise. In Fig 4 (TOP), the graph shows the histogram for 3 single independent RRAM devices sampled at 𝐺𝑡 =50µS and 𝜏=3µS for 300 times. These 3 distributions follow a similar spread centred around 0µS but a randomly different shape, which presents the independent, non-identically distributed challenge of device-to-device variation for completing a theoretical DP guarantee proof. The bottom left plot shows the Q-Q plot as 𝑘 increases from 1 to 3 upon the
Figure 5: (LEFT: RRAM-DP Robustness) Boxplot for the writeverify noise sampling with 3-device representations for a 32×32 crossbar at 20µS, 50µS, 80µS. (MID: RRAM-DP Sampling) Standard deviation (s.t.d.) and program cycles at 20µS, 50µS, 80µS. (RIGHT: RRAM-DP Modeling) S.t.d. and program cycles for 3 random 3-device representations at 50µS. averaged sum of 3𝐶𝑘 devices. It is shown that a 3-device representation aligns nearly perfectly with the diagonal Gaussian under our write-verify sampling method with empirical RRAM stochastic characteristics. The bottom right histogram also expresses the normal Gaussian curve comparison, which says the 3-device representation is a good trade-off between tight privacy evaluation and hardware costs of duplicating 𝑘 RRAM devices. We confirm this Gaussianity condition (𝑘=3) by scaling to 𝜏=1µS, 10µS and 𝐺𝑡 =20µS, 80µS in our empirical qualitative experiments.
4.3
RRAM-DP Relaxation on 𝐺𝑡 & Device Variation
In Fig. 5, we show the stochastic pattern and the correct form of RRAM-DP implementation for a 32×32 1T1R RRAM crossbar. In Fig. 5 (LEFT), a 3-device representation is sampled across the macro at 20µS, 50µS, 80µS upon the convergence of the write-verify algorithm. The box-plot shows consistent Gaussian-like properties at different resistance states, such as symmetry and stable interquartile range spread, over varying 𝜏. The mean offset is also consistent across 20µS, 50µS, 80µS, which is negligible when 𝜏 is small. In Fig. 5 (MID), the e𝑖 ∼Dist(𝜏, 𝐺𝑡 ) is found to have a favourable standard deviation of Δ𝐺 linear relationship with varying 𝜏 bounded by a 2-point distribution. More importantly, this linearity is consistent and independent of e𝑖 ∼Dist(𝜏, 𝐺𝑡 ) can be relaxed all resistance states. This implies Δ𝐺 e to Δ𝐺𝑖 ∼Dist(𝜏) for a simple RRAM-DP sampling that depends only on 𝜏. In Fig. 5 (RIGHT), we demonstrate the hardware lower bound eLB ) for 𝜎 ′ in RRAM-DP mechanism R. We initialize modeling 𝜎 (Δ𝐺 min fixed sampling windows for all devices of 300 empirical error entries at 50µS. Empirical s.t.d. are then estimated over these windows using Chi-Square method [23] (𝛼=0.05) for Gaussian s.t.d. lower bound. We show that individual-level sampling with three examples of 3-device representations follows independent linear curves with slightly different slopes. The linear overlapping region suggests recalibration is only required if any observed window drops below the maintained lower bound. We also report around 10 to 70 expected cycles per device needed for our write-verify algorithm to converge in red color. Cycles may increase exponentially as 𝜏 approaches from 1 to 0.
Count
106
0
-1
-0.5
0
0.5
(k=3) Conductance shift (μS)
Figure 6: (LEFT) Device drift from 15 mins to 12 days. (RIGHT) Final equivalent drift error for 3-device representations. 3.5
25
3.0 2.5 2.0 1.5 1.0 0.5 0
~
MNIST
20 15 10
α=3/10
~
σi (ΔG)≥ σ(ΔGLB) 0 2 4 6 8 10 Stopping Threshold τ
5 0
slope = θ range B α· G range 2ηC 0 2 4 6 8 10 Stopping Threshold τ
MNIST
5 Epsilon ε (δ=10-5)-DP
Cycle Mean
an te e ua r
10 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 Stopping Threshold (μS) Stopping Threshold (μS)
105
σ’min (DPSGD)
ran t
ee
20
0.5
2
Time (s)
nd
1 2 3 4 5 6 7 8 9 10 Stopping Threshold (μS)
20 0.5
104
103
ou
0
(mean offset)
Mean ±1σ
0.85 0.80
4
(k=1)
rb
10
1.0
0.90
30
(k=3)
Mean= -0.002μS σ=0.4499μS
6
0.95
we
20
30 1.0
40
1.00
Lo
HRS
40 1.5
50
8
1.05
~ S.t.d. Value σ(ΔG)
30
1.5
2.0
-g
40
50
DP
2.0
DPgua
50
60
er
S.t.d. Value
60
60 2.5
10
1.10
70
ker
2.5
70 3.0
wea
LRS
3 Devices #1 3 Devices #2 3 Devices #3
ro ng
Conductance (μS)
70
3.0
2-point
80
distribu tion bo und
G t @ 80μS G t @ 50μS G t @ 20μS
st
G t @ 80μS G t @ 50μS G t @ 20μS
K. Wong et al. Normalized Conductance (G/G0)
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
4 3 2
T=
1 0
1 p2
0 2 4 6 8 10 Stopping Threshold τ
Figure 7: (LEFT) Hardware lower bound modeling function. ′ (𝜏) on (MID) Software equivalent lower bound function 𝜎min MNIST dataset. (RIGHT) Final equivalent privacy trade-off function 𝜀 (𝜏) computed at 𝛿=10−5 on MNIST dataset.
4.4
Impact to RRAM-DP from Device Drift
Device drift in RRAM refers to the gradual change in resistance state over time after programming, typically attributed to the evolution of conductive filamentary microstructure [44]. In Fig. 6 (LEFT), we perform an experiment to study the effects of individual device drift after RRAM-DP mechanism under room temperature. Results show that device conductance (𝑘=1) normalised to initial states experienced minimal device drift after 106 s, with mean≈0 and spread≈ ±0.05. In Fig. 6 (RIGHT), we report the final–initial drift error under 3-device representation (𝑘=3) for 106 s. This shows device drift results in an additional Gaussian noise injection with mean≈0µS and standard deviation≈0.45µS, which is equivalent to RRAM-DP modeling at 𝜏=1.5µS from Fig. 5 (RIGHT). This suggests that device drift is introducing an additional one-shot noise offset to RRAM-DP mechanism R at the quasi-saturation level over time. While drift error is small to be considered for one-shot storage mechanism R, it becomes negligible to RRAM-DP-SGD as iteration 𝑇 is large enough for the accumulative program error ≫ drift error. From the privacy perspective, the equivalent per-step drift noise in RRAM-DP-SGD provides negligible additional privacy amplification effect over fixed gradient sensitivity as 𝑇 increases.
4.5
Hybrid Analog-digital Training
To validate device-calibrated DP, we perform hybrid RRAM-DPSGD training on MNIST dataset, where digital update signals are translated into programming pulses to modulate RRAM devices, enabling in-situ (𝑘=3) analog weight adaptation. RRAM-DP Modeling. In Fig. 7 (LEFT), we first model the lineLB ) initialized by Fig. 5 ear hardware lower bound function 𝜎 (Δ𝐺 (RIGHT) at slope 𝛼=3/10. At runtime, the sampling window remove oldest error sample and append new final read-out error at each
RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning
20 0
0
5
10 15
20 25
Iterations T
30
22000
-1
21000
-2
0
5
10 15
20 25
Iterations T
30
-3
training iteration. In Fig. 7 (MID), we transform the hardware lower ′ bound into equivalent software lower bound function 𝜎min that is scaled by MNIST dataset sensitivity, such as batch size and learning rate, for hardware 𝜏 as a control variable. In Fig. 7 (RIGHT), we fur′ ther compute the (𝜀, 𝛿)-DP form by transforming 𝜎min into 𝜇𝑟 -GDP 2 using 𝑇 = 1/𝑝 as the correct convergence scaling from RRAM-DPSGD in theorem 3.2 and (𝜀, 𝛿)-DP conversion in theorem 2.1. Thus, by choosing any write-verify control variable 𝜏, we get the desirable RRAM-DP-SGD privacy level calibrated for MNIST. RRAM-DP-SGD Training. In Fig. 8 (LEFT), we perform hybrid analog-digital training with RRAM-DP-SGD and compare the actual hardware noise with simulation noise, DP-SGD training, and normal training using a clipping-free and noise-free optimizer. While the normal training is capped at ∼90% accuracy, we train privately using noise calibrated at 𝜀=0.1. First, DP-SGD uses the standard 𝜇-GDP accountant [13] to estimate isotropic Gaussian noise injection. Second, RRAM-DP-SGD Hardware uses 𝜏 ≈10 from Fig. 7 (RIGHT) to inject realistic RRAM noise at 𝜀=0.1. Third, RRAM-DP-SGD Simulation uses 𝜏=10 with empirical s.t.d. collected for all devices in Fig. 5 (RIGHT) to inject diagonal Gaussian noise with a fixed s.t.d. lookup table. Results show that RRAM-DP-SGD simulation closely mimics the RRAM-DP-SGD hardware training behaviour, both achieving accuracy at ∼80%. This also shows 𝑘=3 closely approximates the theoretical Gaussian noise. Meanwhile, DP-SGD achieves similar accuracy of ∼82% at the same theoretical privacy level, which proves the tight hardware estimation with 𝜏 ≈10. Training Program Cycles. In Fig. 8 (RIGHT), we record the total programming cycles of RRAM-DP-SGD along the hardware training iterations. Given 760 weights in the MLP network, roughly 21000– 23000 cycles are performed in each iteration. Consider the 3-device representation, we randomly sample 3 devices for the same weight to program, where each device is expected to program 23000/(760*3)≈10.1 times. We find this result closely aligns with program cycles shown in RRAM-DP Sampling in Fig. 5 when 𝜏=10.
Privacy Protection against MIAs
In Fig. 9, we identify the privacy protection power between normal training without privacy-preserving techniques, DP-SGD, and RRAM-DP-SGD hardware all trained in Fig. 8. For the MIA attack, we use Likelihood Ratio Attack (LiRA) [7] with 256 shadow models (N=256) performed under 10,000 balanced attack samples to capture membership and non-membership training distribution on the resized 8×8 MNIST dataset. The effectiveness of the LiRA attack can be evaluated from two perspectives: (1) an area under the ROC curve (AUC) close to 0.5, indicating near-random discrimination capability;
-4
Unbiased Detection RRAM-DP-SGD; AUC=0.4975 DP-SGD; AUC=0.4972
-1
-2
10
10
20500
0
10
10
21500
Figure 8: Experimental results for RRAM-DP-SGD hardware training (𝜀=0.1) as compared to DP-SGD, normal training.
4.6
10
Unbiased Detection Normal Training; AUC=0.5348
10
22500
20000
0
True Positive Rate
RRAM-DP-SGD Hardware RRAM-DP-SGD Simulation DP-SGD Software Normal Training
40
23000
True Positive Rate
60
23500
-3
10
10 -4 10
-3
10
-2
10
-1
10
False Positive Rate
10
0
-4
10 -4 10
-3
10
-2
-1
10
10
False Positive Rate
10
0
Figure 9: ROC curves of MIA (LiRA (N=256)) against normal training, DP-SGD, and RRAM-DP-SGD hardware training calibrated at (𝜀=0.1, 𝛿=10−5 )-DP protection level. 140 120 100
σ’min (DPSGD)
80
10
RRAM-DP-SGD Hardware
80 60 40
CIFAR10 CIFAR100 STS-B SST-2
slope = θ range B α· G range 2ηC
20 0 0 1 2 3 4 5 6 7 8 9 10 Stopping Threshold τ
Epsilon ε (δ=O(1/n))-DP
24000 Total Programming Cycles
Test Accuracy (%)
100
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
CIFAR10 CIFAR100 STS-B SST-2
5 4 3 2 1 0
T=
1 p2
0 1 2 3 4 5 6 7 8 9 10 Stopping Threshold τ
Figure 10: (LEFT) Software equivalent lower bound function ′ (𝜏) on vision/language datasets. (RIGHT) Final equivalent 𝜎min privacy trade-off function 𝜀 (𝜏) computed at 𝛿=O(1/n) on vision/language datasets. and (2) a low true positive rate (TPR) at low false positive rates (FPR), reflecting limited statistical power when the significance level is small. As shown in Fig. 9, normal training exhibits mild privacy leakage (AUC = 0.5348). In contrast, both DP-SGD and RRAM-DP-SGD achieve AUC values close to 0.5 and suppressed TPR in the lowFPR regime. This confirms the strong protection power of utilizing device-realistic noise for DP private learning.
4.7
Unlocking CiM-based Training with Privacy and Pretraining Techniques
RRAM-DP Modeling. To fully utilize noise of in situ CiM-training, we scale up RRAM-DP-SGD training from MNIST to CIFAR10/100 (vision) and STS-B/SST-2 (language) tasks. Similarly, we identify the eLB ) = 𝛼 ·𝜏 with slope 𝛼 = 3/10 same linear lower bound model of 𝜎 (Δ𝐺 ′ derived from Fig. 5 to estimate 𝜎min for computing the composition ′ of 𝜇𝑟 by theorem 3.2. Here, a steeper slope of lower bound 𝜎min refers to a stricter privacy guarantee at the same hardware noise injection as dataset-specific training hyper-parameters scale the gradient sensitivity. Then, we formulate 𝜀 (𝛿 follows Table 1) against 𝜏 relationship in Fig. 10 (RIGHT) to calibrate target privacy (𝜀, 𝛿)-DP at some 𝜏 for different dataset training configurations. It is observed that the flatter slope in Fig. 10 (LEFT) has easier privacy control in Fig. 10 (RIGHT) with RRAM-DP-SGD over four datasets. Pretraining Effects. Usually, a larger dataset length or a dataset with more classes, such as CIFAR100, requires a richer visual pretraining dataset and larger batch size to improve utility, where the ′ later one causes the steep lower bound 𝜎min slope to be more difficultly compared at the same 𝜀 across datasets. Here, we adopt the
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
ε=0.5
ε=1
ε=2
ε=3
CIFAR10 (δ=10-5 )
ε=5
60 40 20 ε=0.5
ε=1
ε=2
ε=3 -4
STS-B (δ=10 )
20 0
100
80 Non-DP: 85.8 [12]
0
40
ε=5
80
ε=0.5
ε=1
ε=2
ε=3
CIFAR100 (δ=10-5 )
ε=5
30 20
40 20 0
ε=0.5
ε=1
ε=2
ε=3 -5
SST-2 (δ=10 )
ε=5
same convergence scaling of 𝑇 =1/𝑝 2 from theorem 3.2 by fixing the batch size to compute the iterations 𝑇 needed for training and replacing batch normalization layers with group normalization for per-sample clipping. Upon these setups, we simulate the CiM noise in the same way from Fig.5 and Fig.8 for privately training CIFAR10, CIFAR100, STS-B, and SST-2 over the deployed pretraining model in Fig. 11. When 𝜀 = 5, their accuracy is equal to 89.6%, 71.8%, 83.4%, and 92.0%, which is comparable to their non-private baseline training from scratch [43] and pretrained model [12, 27], and better than state-of-the-art DP-SGD training from scratch [11, 34, 35]. This suggests pretraining technique is key to high utility CiM-based training with better tolerance against noise accumulation.
High-Performance RRAM-DP-SGD
Energy Saving. In Fig. 12, energy savings are estimated by normalizing A100 energy consumption under theoretical limit in sec. 4.1 setting. The system operates at a 500MHz clock frequency. For in-memory computing, 512 RRAM crossbar tiles (32×32 arrays) execute 512×32×32 parallel MAC operations, with energy per MAC estimated at 0.447pJ (≈FLOPS/2), including peripheral contributions (ADC: 72.3fJ, DAC: 46.3fJ, op-amp: 0.51pJ, RRAM driver: 0.27pJ, MUX: 0.12pJ)[5, 10, 25]. RRAM inference uses a 0.2V read voltage with 50µS average conductance, while programming employs SET/RESET voltages of 1.2V/2.2V, with row-parallel programming 512×32 devices executing at 29.2pJ per device per cycle. For the RRAM-DP-SGD accelerator at (𝜀=2, 𝛿=O(1/n))-DP, we achieve 54×-57× energy reduction against A100 and 3×-3.2× against DiVa-GEMM [32]. Despite RRAM write overheads, the main energy cost arises from per-sample gradient batch computations, with smaller batches benefiting less from energy savings. Time Speedup. In Fig. 13, speedup is similarly estimated by normalizing A100 computation time under theoretical limit in sec. 4.1 setting. The system leverages massive MAC parallelism through in-memory computing with 512×32×32 operations per MAC step and row-parallel programming supporting 512×32 devices in one operation. Each programming cycle takes 40 ns, incorporating the time for SET/RESET and readout operations under the specified programming scheme. Each MAC operation takes 20 ns, with the overall execution highly optimized by the 500 MHz clock and parallel tile architecture. We achieve 2.5×-2.7× speedup compared to A100 and 1.7×-1.8× compared to DiVa-GEMM [32] at (𝜀=2, 𝛿=O(1/n))-DP, with major time consumed in MAC compared to RRAM program time. This demonstrates the promising in-memory computing paradigm with privacy guarantees.
~56X
A100 (40GB) DiVa-GEMM [32] RRAM-DP-SGD
40
0
60
Figure 11: Experimental results for RRAM-DP-SGD with pretraining techniques (𝜀=2) as compared to non-private training (Non-DP) and DP-SGD from scratch.
4.8
50
10
Non-DP: 94.8 [27]
~57X
~18X
~57X
~54X
~18X
0.8
~3.2X
~3.0X
1.0
RRAM-DP SGD Energy
0
60
60
1%
2%
6%
1%
99% 98% 94% 99%
0.6
~18X
~18X
0.4 Write Energy MAC Energy
0.2
1X
1X
1X
1X
CIFAR10
CIFAR100 STS-B(r=2) SST-2(r=4) (ε=2, δ=O(1/n))-DP
0.0
C’10 C’100 STS’ SST’ (ε=2, δ=O(1/n))-DP
Figure 12: (LEFT) Experimental results for hardware energy savings as compared to GPU (A100) and DiVa-GEMM. (RIGHT) Energy consumption for RRAM. 3.2 ~2.7X
2.8
~2.6X
2.4
~1.7X
2.0 1.6
~1.5X
1.2
1X
0.8 0.4 0.0
~2.7X
~2.5X ~1.8X
~1.5X
~1.5X
~1.5X
1X
1X
1X
A100 (40GB) DiVa-GEMM [32] RRAM-DP-SGD CIFAR10
CIFAR100 STS-B(r=2) SST-2(r=4) (ε=2, δ=O(1/n))-DP
1.0 RRAM-DP SGD Time
Tang et al. [34] De et al. [11] Tramèr & Boneh [35]
20
70
Non-DP: 79.5 [43]
Energy Saving
40
80
Speedup
60
100
Test Accuracy(%)
100
Non-DP: 94.8 [43]
Test Accuracy(%)
80
Test Accuracy(%)
Test Accuracy(%)
100
K. Wong et al.
0.8 0.6
1%
2%
6%
1%
99% 98% 94% 99%
0.4 0.2 0.0
Write Time MAC Time C’10 C’100 STS’ SST’ (ε=2, δ=O(1/n))-DP
Figure 13: (LEFT) Experimental results for hardware speedup as compared to GPU (A100) and DiVa-GEMM. (RIGHT) Time consumption for RRAM. Discussion on SOTA privacy accelerators. DiVa [32] is the first architecture design dedicated to DP training, replacing the conventional weight-stationary systolic GEMM array with an outer-product dataflow engine to better accommodate the irregular, small innerdimension matrix multiplications associated with per-sample gradient computation. However, DiVa remains a fully digital implementation with conventional memory hierarchies for MAC operations. Although cryptographically safe pseudo-random number generator (CSPRNG) addresses floating-point error for digital accelerators, it renders additional latency and energy overhead (see secure mode in Opacus [42] and Jax-privacy [28]).
5
Conclusion
Former works assume idealized noise generation in hardware for privacy applications, which is not always true in real device statistics. In this work, we propose RRAM-DP co-design to connect real RRAM device randomness for private on-device data storage and in-memory SGD with formal (𝜀, 𝛿)-DP privacy definition, with theoretical proof under 𝜇-GDP analysis framework. Our design identifies a robust RRAM noise sampling method with a single control variable, validated through empirical device statistics to inject sufficient RRAM-DP mechanism noise for privacy protection. Additionally, we show that the pretraining technique inspired by DP-SGD can mitigate CiM-induced noisy learning, enabling near-lossless accurate CiM-training as compared to normal training. Our RRAM-DP-SGD accelerator establishes a novel computing paradigm for both efficient and private AIoT edge learning.
6
Acknowledgement
This research was conducted by ACCESS – AI Chip Center for Emerging Smart Systems, supported by the InnoHK initiative of the Innovation and Technology Commission of the Hong Kong Special Administrative Region Government; supported by Hong Kong Research Grant Council - General Research Fund Scheme (Grant No. 17202422, 17212923, 17215025) Theme-based Research (Grant No.T45701/22-R), and Strategic Topics Grant (Grant No.STG3/E-605/25-N).
RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning
References [1] Martin Abadi et al. 2016. Deep Learning with Differential Privacy. In Proceedings of Conference on Computer and Communications Security (CCS ’16). [2] Maya Anderson et al. 2025. Is My Data in Your Retrieval Database? Membership Inference Attacks Against Retrieval Augmented Generation. In Proceedings of the International Conference on Information Systems Security and Privacy (ICISSP ’25). [3] Zhiqi Bu et al. 2020. Deep Learning With Gaussian Differential Privacy. Harvard Data Science Review (2020). [4] Zhiqi Bu et al. 2023. Differentially Private Optimization on Large Model at Small Cost. In Proceedings of the 40th International Conference on Machine Learning (ICML ’23). [5] Fuxi Cai et al. 2020. Power-efficient combinatorial optimization using intrinsic noise in memristor Hopfield neural networks. Nature Electronics (2020). [6] Nicholas Carlini et al. 2021. Extracting Training Data from Large Language Models. In Proceedings of the USENIX Conference on Security Symposium (SEC ’21). [7] Nicholas Carlini et al. 2022. Membership Inference Attacks From First Principles. In Proceedings of the IEEE Symposium on Security and Privacy (SP ’22). [8] Nicholas Carlini et al. 2023. Extracting Training Data from Diffusion Models. In Proceedings of the USENIX Conference on Security Symposium (SEC ’23). [9] Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter. 2017. A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets. arXiv preprint 1707.08819 (2017). [10] Xiaofeng Chu, Can Liang, and Zeyu Cai. 2025. A 67.51 dB SNDR 137.8 𝜇 W SARAssisted Slope ADC with High-Linearity Analog Slope and Low Switching Energy. In Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS ’25). [11] Soham De et al. 2022. Unlocking High-Accuracy Differentially Private Image Classification through Scale. In Proceedings of the ICML Workshop on Theory and Practice of Differential Privacy. [12] Jacob Devlin et al. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL ’19). [13] Jinshuo Dong et al. 2022. Gaussian Differential Privacy. Journal of the Royal Statistical Society Series B: Statistical Methodology (2022). [14] Cynthia Dwork et al. 2006. Calibrating Noise to Sensitivity in Private Data Analysis. In Proceedings of Conference on Theory of Cryptography (TCC ’06). [15] Cynthia Dwork and Aaron Roth. 2014. The Algorithmic Foundations of Differential Privacy. Foundations and Trends in Theoretical Computer Science (2014). [16] Joseph P. Romano E.L. Lehmann. 2022. Testing Statistical Hypotheses. Springer Texts in Statistics (2022). [17] Matt Fredrikson et al. 2015. Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS ’15). [18] Jingyan Fu et al. 2021. Memristor-Based Variation-Enabled Differentially Private Learning Systems for Edge Computing in IoT. IEEE Internet of Things Journal (2021). [19] Justin Hsu et al. 2014. Differential Privacy: An Economic Method for Choosing Epsilon. In Proceedings of the IEEE 27th Computer Security Foundations Symposium (CSF ’14). [20] Daniele Ielmini and Giacomo Pedretti. 2025. Resistive switching random-access memory (RRAM): Applications and requirements for memory and computing. Chemical Reviews (2025). [21] Jiankai Jin et al. 2022. Are We There Yet? Timing and Floating-Point Attacks on Differential Privacy Systems. In Proceedings of the IEEE Symposium on Security and Privacy (SP ’22). [22] Peter Kairouz et al. 2021. Practical and Private (Deep) Learning Without Sampling or Shuffling. In Proceedings of the 38th International Conference on Machine Learning (ICML ’21). [23] Achim Klenke. 2014. Probability Theory: A Comprehensive Course. Universitext (2014). [24] Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. 2009. CIFAR-10 and CIFAR-100. http://www.cs.toronto.edu/∼kriz/cifar.html. [25] Sanghyun Lee and Youngmin Kim. 2023. An energy efficient 7.59-ENOB 50 MS/s flash-SAR ADC in 65-nm CMOS. In Proceedings of the IEEE International Midwest Symposium on Circuits and Systems (MWSCAS ’23). [26] Jinchang Liu et al. 2025. Error-aware probabilistic training for memristive neural networks. Nature Communications (2025). [27] Yinhan Liu et al. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint 1907.11692 (2019). [28] Ryan McKenna et al. 2026. JAX-Privacy: A library for differentially private machine learning. arXiv preprint 2602.17861 (2026). [29] Valerio Milo et al. 2021. Accurate Program/Verify Schemes of Resistive Switching Memory (RRAM) for In-Memory Neural Network Circuits. IEEE Transactions on Electron Devices (2021). [30] Ilya Mironov. 2012. On Significance of the Least Significant Bits for Differential Privacy. In Proceedings of the ACM Conference on Computer and Communications Security (CCS ’12). [31] Arvind Narayanan and Vitaly Shmatikov. 2008. Robust De-anonymization of Large Sparse Datasets. In Proceedings of the IEEE Symposium on Security and Privacy (SP
ICCAD’2026, November 8–12, 2026, San Jose, CA, USA
’08). [32] Beomsik Park et al. 2022. DiVa: An Accelerator for Differentially Private Machine Learning. In Proceedings of the IEEE/ACM International Symposium on Microarchitecture (MICRO ’22). [33] Haoxiong Ren et al. 2025. When Pipelined In-Memory Accelerators Meet Spiking Direct Feedback Alignment: A Co-Design for Neuromorphic Edge Computing. In Proceedings of the IEEE/ACM International Conference on Computer-Aided Design (ICCAD ’25). [34] Xinyu Tang et al. 2023. Differentially Private Image Classification by Learning Priors from Random Processes. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS ’23). [35] Florian Tramèr and Dan Boneh. 2021. Differentially Private Learning Needs Better Features (or Much More Data). In Proceedings of the International Conference on Learning Representations (ICLR ’21). [36] Alex Wang et al. 2018. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Proceedings of the EMNLP Workshop on BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. [37] Ning Wang et al. 2019. Collecting and Analyzing Multidimensional Data with Local Differential Privacy. In Proceedings of the IEEE International Conference on Data Engineering (ICDE ’19). [38] Zhongrui Wang et al. 2019. In situ training of feed-forward and recurrent convolutional memristor networks. Nature Machine Intelligence (2019). [39] Zhongrui Wang et al. 2020. Resistive switching materials for information processing. Nature reviews materials (2020). [40] Kwunhang Wong et al. 2025. SNNGX: Securing Spiking Neural Networks with Genetic XOR Encryption on RRAM-based Neuromorphic Accelerator. In Proceedings of the IEEE/ACM International Conference on Computer-Aided Design (ICCAD ’24). [41] Xiaoju Ye. 2023. Calflops: a FLOPs and Params calculate tool for neural networks in pytorch framework. https://github.com/MrYxJ/calculate-flops.pytorch. [42] Ashkan Yousefpour et al. 2021. Opacus: User-Friendly Differential Privacy Library in PyTorch. In Proceedings of the NeurIPS Workshop on Privacy in Machine Learning. [43] Sergey Zagoruyko and Nikos Komodakis. 2017. Wide Residual Networks. arXiv preprint 1605.07146 (2017). [44] Furqan Zahoor, Tun Zainal Azni Zulkifli, and Farooq Ahmad Khanday. 2020. Resistive random access memory (RRAM): An overview of materials, switching mechanism, performance, multilevel cell (MLC) storage, modeling, and applications. Nanoscale research letters (2020). [45] Fengjiao Zhang, Zhao Pan, and Yaobin Lu. 2023. AIoT-enabled smart surveillance for personal data digitalization: Contextual personalization-privacy paradox in smart home. Information & Management (2023).