Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch Heeeon Lee
Yonsei University Republic of Korea [email protected]
arXiv:2609.34657v1 [cs.AR] 28 Sep 2026
Hyunmo Sung
Hyunwoo Nam
Yonsei University Republic of Korea [email protected]
Jay Hwan Lee
Yeonsoo Kim
Yonsei University Republic of Korea [email protected]
Yonsei University Republic of Korea [email protected]
Yonsei University Republic of Korea [email protected]
Seongho Jeong
Shinhyung Yang
Bernd Burgstaller
Yonsei University Republic of Korea [email protected]
Kiel University Kiel, Germany [email protected]
Abstract Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. However, PyTorch and other DL frameworks lack compiler support for making this decision on the code they lower: existing offloading frameworks target hand-written C/C++ programs, while those that address DL fix the candidate set to a list of operator types before lowering. We present Torch-PIM, a compiler framework that uses profile-guided optimization (PGO) to decide host-versus-PIM placement over the loop nests that progressive lowering materializes. Every parallel loop nest the pipeline emits enters the candidate space, and each is assessed in two stages: the amount of work it carries, and its memory boundedness. Every quantity the assessment consumes is profiled on the host or obtained from the multi-level intermediate representation (MLIR) of the code. Across PIM configurations of 32 to 128 cores, Torch-PIM’s offloading decisions yield speedups of up to 8.6 × on tensor operators, 2.9 × on MLP, 4.4 × on Attention, 5.1 × on GPT-J6B, and 3.6 × on LLaMA-7B over CPU-only execution.
1
Junyong Heo
Yonsei University Republic of Korea [email protected]
Introduction
Modern DL workloads are increasingly limited by data movement rather than arithmetic throughput [9, 10, 43]. PIM targets this bottleneck by integrating compute units directly into or near memory. However, PIM introduces a software challenge: deciding which code regions to offload from the host to PIM [2]. Incorrect offloading decisions are costly; naive policies based on last-level cache misses per kilo instructions (MPKI) can incur up to 41.6× slowdowns because the context-switch and data-dependency overhead incurred at each transition between the host and PIM are not accounted
Yonsei University Republic of Korea [email protected]
for [39], and manual code offloading is entirely impractical for programmers [20]. Automated compiler-driven offloading is therefore necessary. Prior work has demonstrated the advantage of offloading for PIM-integrated systems. However, designing an effective compiler-driven DL offloading framework for such systems faces three challenges. First, no framework automates a PIM offloading decision for code that is produced from the DL framework after lowering. Only hand-written C/C++ benchmarks, primarily graph workloads, have been the main target of automatic offloading [1, 5, 8, 11, 12, 17, 29, 39–41]. Although offloading methods for DL operators have been investigated, they restrict the target operators to predetermined candidates and focus on orchestrating the fixed operators based on their strategies, such as subgraph patterns [19, 34], or optimizing the performance on PIM [6, 25, 35]. Second, the granularity of offloadable regions does not fit the structure of DL workloads. DL workloads are expressed as tensor operators, and a single operator commonly has several phases with different behavior, while a model chains hundreds of such operators [42]. At the operator level, computeand memory-bound loop nests are grouped into one unit: more than half of the candidate layers in a CNN can be placed neither wholly on the host nor wholly on PIM [34], and offloading such an operator as a whole can result in worse performance than not offloading it [5, 17]. Offloading at the source loop level is impractical, since a tensor operator in DL frameworks carries no loop for a programmer to annotate or an instrumenter to mark [24, 41]. The basic block level risks frequent CPU–PIM switches: if consecutive blocks are assigned to different targets, every boundary incurs data-transfer and synchronization cost. Third, the profiling cost for offloading decisions is expensive. Prior work executes entire regions on PIM to decide the best offloading [39]; executing compute-bound regions
Lee et al.
on PIM is redundant and expensive. Another work profiles each input separately, which costs up to 11.6 × the execution time of the workload it optimizes [5]. We present Torch-PIM, a compiler framework that makes offloading decisions over loop nests produced after lowering from PyTorch [31]. Torch-PIM lowers a PyTorch model through Torch-MLIR [26] to linalg, then to explicit loop nests, and admits each parallel loop nest the pipeline produces as an offloading candidate. Nothing designates these candidates: they are not a list of operator types, not patterns identified in advance, and not regions a programmer annotated. Whatever the lowering emits enters the candidate space. One operator therefore yields several candidates, each judged on its own, so an operator that mixes memorybound and compute-bound loop nests is split along that line instead of being placed as a unit. Torch-PIM makes each decision from host-side information alone. For every loop nest it gathers three quantities: the elapsed time of the CPUonly execution; the achieved DRAM read bandwidth and the store-bound cycle fraction, derived from five native hardware performance-monitoring events collected through PAPI [14] in a single counter set; and the volume of data a PIM placement would transfer, obtained from the MLIR. These are weighed against a switch overhead calibrated once on the target platform. No candidate is ever executed or simulated on the PIM device. We evaluate Torch-PIM on 15 PyTorch tensor operators, 2 layers of tensor operators, and 2 real-world benchmarks, LLaMA-7B [36], and GPT-J-6B [37]. Torch-PIM’s decisions yield speedups of up to 8.6 × on tensor operators, 4.4 × on layers of tensor operators, 5.1 × on GPT-J-6B, and 3.6 × on LLaMA-7B over CPU-only execution. In summary, Torch-PIM makes the following contributions: • Automated host-vs-PIM offloading for the code produced from DL framework after lowering. Torch-PIM lowers PyTorch models through MLIR and makes offloading decisions autonomously over the code it generates. • Compiler-materialized loop nests as the decision unit for offloading. Torch-PIM identifies offloading candidates at the loop nests that progressive lowering materializes, which do not exist until the compiler produces them and therefore do not require an offloading decision at the abstraction level of the input source code. Thus, Torch-PIM eliminates the abstraction mismatch between abstract source-level ML operators and their lowered representations, which oftentimes contain memory- and compute-bound loops that require individual offloading decisions. • Profile-guided offloading decision with microarchitectural events. Torch-PIM employs profiling information obtained from micro-architectural events
of the underlying host architecture to identify memorybound loops: elapsed CPU cycles, DRAM read bandwidth, and the fraction of cycles a loop spends bound on stores. These are weighed against the overhead of offloading, calibrated once on the target platform. • An offload rule built entirely from host-side measurements. Torch-PIM reaches an offloading decision without any measurement of PIM-side behavior. • Evaluation on real PyTorch framework. We conduct an extensive experimental evaluation on code generated from PyTorch framework, consisting of 15 PyTorch tensor operators, 2 layers of tensor operators, and 2 real-world benchmarks comprising LLaMA7B [36], GPT-J-6B [37]. PIM-offloading of Torch-PIM achieves speedups of up to 8.6 × on 32, 64, 128 PIM cores over CPU-only execution.
2
Background
2.1
Modern DL Workloads and the Viability of PIM
Data movement limits performance across a wide range of workloads. A top-down analysis of 77K functions drawn from 345 applications identified 144 functions that account for at least 3% of total cycles and spend more than 30% of their execution bound on memory; among these, the functions whose bottleneck lies in DRAM bandwidth benefit most from computation placed near memory [30]. The same behavior appears in DL inference. A layer-level analysis of neural network models for edge devices attributes a large fraction of inference time to data movement [3], and roofline analyses of large language models establish that the autoregressive decode phase operates at low arithmetic intensity and is limited by memory bandwidth [44]. However, identifying the best execution unit solely from workload names or the operator identities is difficult due to the different characteristics of the host architectures. FFN kernels that are compute-bound on datacenter GPUs move into the memory-bound region on edge devices, where cache capacity and LPDDR bandwidth are limited [38]. By addressing the memory bottleneck, PIM has attracted attention in academia, and has moved from academic proposals into commercial silicon. HBM2 [22, 32] and GDDR6 [21] were followed in 2026 by an LPDDR5X-based device [13]. Standardization is also underway: following the JESD209-6 LPDDR6 standard [15], a LPDDR6 Processing-in-Memory standard is in development [16]. These devices are designed to minimize changes to the host memory system, so the same physical memory is reached both through the host’s ordinary access path and through the PIM compute path. Execution on PIM differs from execution on the host in three respects. First, compute units placed inside the banks operate under area and power budgets and are simpler than host cores. The compute unit of a commercial PIM device is an in-order 32-bit RISC core running at 350 MHz that
Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch
requires eleven hardware threads to fill its pipeline. Second, switching execution between the host and PIM carries a cost. Data to be offloaded may reside in the host caches and must be flushed at the switch point, and a transfer in the opposite direction is required when the host reads the result. Third, the host and PIM cannot write the same data concurrently during an offloaded region. In the commercial LPDDR5X PIM device, the host must place the device back in conventional DRAM mode before reading the result [13]. 2.2
Torch-MLIR
Torch-MLIR [26] maps a PyTorch model into MLIR. It traces the model into a graph of tensor operations and emits the torch dialect, which mirrors the PyTorch operator set with static types attached. Once the model is in MLIR, it can be lowered through the dialects MLIR provides and the passes at each level can be varied, so a compiler for a new target can be built by choosing where to intervene rather than by writing a backend from scratch.
3
Related Work
3.1 Existing Automatic PIM Offloading Frameworks Automatic PIM offloading frameworks differ in the granularity at which they form candidate regions, from individual instructions to entire kernels, and in the information their decision rules consume. In all of them, however, the candidate set is fixed by boundaries of the source program or the framework interface already defines. At the finest granularity, PEI [1] defines a set of memory-side primitives and dispatches individual operations to them according to their locality. GraphPIM [29] maps atomic instructions in graph workloads onto memory-side execution. PIMProf [39] partitions a program into basic blocks and functions, simulates each on both sides to obtain per-unit costs, and solves for a partition that accounts for the data movement induced at the boundaries. A3 PIM [17] uses the same granularity as PIMProf, and classifies functions from a static analysis of their access patterns, and RDPIM [40] derives partitions for graph applications from the connectivity of the input graph. CoPIM [41] decides over programmer-written loop bodies, sampling a fraction of the iterations at runtime to estimate the cache behavior of the whole; LooPIM [24] and other loop-oriented approaches [27, 28] follow a similar path. The unit these frameworks decide over is, in every case, one whose boundaries the source determines. The compiler recovers instructions and basic blocks, but their number and extent follow from the control flow the programmer wrote; functions and loop bodies are written in the source directly; operators are fixed by the framework interface. No such correspondence holds for a compiled DL model. A tensor operator specifies an iteration space and nothing more; the pipeline’s fusion and decomposition decisions determine how many loop nests execute it, in what order, and where one
Stage 1: Extraction and profiling
PyTorch
MLIR lowering
Loop decision (CPU or PIM)
Loop extract (parallel loops)
Profile (LLVM inst. + PMU)
Reinsert (offload boundaries)
CPU–PIM binary
Stage 2: Decision and reinsertion
Figure 1. Torch-PIM overview. Stage 1 lowers PyTorch code to MLIR, extracts parallel loop candidates, and profiles them using LLVM instrumentation and PMU events. Stage 2 assigns loop-level CPU/PIM decisions and reinserts offload boundaries into code generation. ends and the next begins. The evidence these frameworks collect is likewise unavailable at compile time: a static analysis of access patterns alone does not establish that a region is bound by memory on the machine that runs it [17], sampling a subset of iterations characterizes only the iterations sampled [41], and simulating both sides of the execution costs more than a compilation pipeline can absorb [39]. 3.2
PIM Compilers for DL Frameworks
Several compiler infrastructures provide a pipeline for DL frameworks targeting PIM, yet they universally assume the offloading decision is already made, and optimize resource distribution on PIM cores. OptiPIM [25] lowers PyTorch through Torch-MLIR down to linalg and affine loop nests. It formulates an integer linear programming (ILP) problem to map tiling, data layouts, and indexing onto PIM resources, supposing that the offloading decision has already been made. Other frameworks make the same assumption. ATiM [33] extends TVM to auto-generate kernels and autotune intra- and inter-DPU optimization spaces for UPMEM, but operates on operations already designated for offloading. CINM [19] provides an MLIR dialect layer for compute-inmemory and compute-near-memory devices, but it doesn’t consider “whether to offload”. Likewise, other end-to-end DNN compilers [35] and programming frameworks [6, 18] automate data distribution and kernel execution (how to offload), yet leave the decision of what to offload entirely to the programmer. Torch-PIM fills this gap by autonomously deciding CPU-PIM placement across the loop nests it lowers from the DL framework.
4
Torch-PIM Design
4.1
Overview of Torch-PIM
Torch-PIM is a static-shape DL compiler for improving performance on CPU-PIM integrated systems. Torch-PIM operates in two stages. The first stage lowers operations and characterizes each candidate loop nest both statically and dynamically: static analysis at the MLIR level, and profiling
Lee et al.
at the LLVM IR level with hardware performance counters. The second stage determines the CPU–PIM execution boundary: Torch-PIM applies the offloading rule to the collected characteristics and emits offload markers in the generated program. 4.2
Progressive Lowering in MLIR
Torch-PIM progressively lowers a PyTorch model through MLIR dialects, identifying candidate loop nests and their static characteristics along the way. The remaining inputs are measured by profiling the instrumented program on the host. torch: obtains input shapes. A PyTorch model is first lowered to the torch dialect by Torch-MLIR [26]. To do so, Torch-PIM requires every tensor extent to be a compile-time constant, which holds when a model is exported for a fixed input shape. The same constant later appears as a dimension of a memref type, as a loop bound materialized from it, and as an allocation size in the emitted code. 1 2
#map = affine_map<(d0, d1) -> (d0, d1)> #par = ["parallel", "parallel"]
3 4 5 6 7 8 9 10
{indexing_maps = [#map, #map], iterator_types = #par} ins (% outs(% ^bb0(% linalg.yield % } -> tensor<64x64xf32>
Figure 2. One loop nest at the linalg dialect. The nest states its read set in ins and its write set in outs, and both operands carry a constant shape, so the bytes a placement would have to move follow from the operand types alone: 64 × 64 × 4 = 16,384 on each side. linalg: extracts loop-nest boundaries, logical operands. Torch-PIM first identifies loop nests at the linalg dialect, where the boundaries of its decision units are settled. Fusion and decomposition decisions act on structured operations, so how many loop nests are extracted and where their boundaries fall is determined at this level: a softmax becomes eight loop nests and an attention layer fourteen. No function of the input program corresponds to any of them. Torch-PIM also determines at this level how many bytes a loop nest would have to move if it were placed on PIM. Each of these nests appear here as a single linalg operation that declares what it reads and what it writes as whole tensors, through its ins and outs operands like the example depicted in Fig. 2. The read and write set of a nest are therefore available as logical units with a static extent, before lowering scatters them into individual memory accesses. The
size of each operand follows from the product of its constant shape dimensions times the element width. Torch-PIM then bufferizes these operations. Immediately after bufferization, once memref types are settled and before following passes erase them, Torch-PIM records two things for each operand: whether it is a memref, and the allocation it resolves to. Operands that are not memref values are charged no transfer. Neither is available before this point, because linalg operands are values, not storage. The ins and outs of a nest name the tensors it reads and writes, from which the sizes are known, but a tensor at this level has no address: two operands holding the same tensor may end up sharing one buffer or occupying two, and an operation that appears to produce a new tensor may in fact overwrite its input. Bufferization decides the placement, assigning a buffer only to the values that need one and leaving the rest in place. It settles which data would actually cross a host–PIM boundary. Lowering to the LLVM dialect afterward replaces every memref operand with a bare pointer and address arithmetic, and the information disappears. Several operands can name the same buffer. An output written in place carries the buffer of the input it overwrites, and a value threaded through successive operations reappears as an operand of each. Identifying operands by their allocation, a buffer that is read and then written in place counts as one transfer instead of two, and a buffer that a producer writes and a consumer within the same segment reads back is counted once. scf: counts iteration. Torch-PIM reads the scf dialect and counts iterations of each loop nest. linalg declares which dimensions of a nest are parallel and which are reductions, but that is the dependence structure of the iteration space, not its extent. Bufferization materializes that space into actual loops: parallel dimensions become scf.parallel, reduction dimensions scf.for with scf.reduce. The two kinds of iteration are now distinct operations, each carrying constant bounds. Only the loops materialized as scf.parallel are passed to the offloading rule, and the bounds and stride of each give its iteration count, e.g., 0 to 64 step 1 yields 64. From these, Torch-PIM collects three values: the product of the iteration counts over the parallel dimensions alone, the product over every dimension of the nest, and the trip count of the enclosing loops. The first two feed the stageone conditions of Section 4.4; the third, together with the buffer sizes of the previous step, feeds the data-movement calculation of Section 5.1. 4.3
Profile-Guided Optimization
Torch-PIM obtains the inputs to its offloading decision by instrumenting the lowered LLVM IR with marker calls around each candidate loop nest and profiling it on the host. The markers are thin wrappers built on PAPI [14]. At the entry
Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch
Table 1. Native events programmed as a single set on the profiling host. Event
Term
offcore_requests.all_data_rd EXE_ACTIVITY.BOUND_ON_STORES CYCLE_ACTIVITY.STALLS_TOTAL EXE_ACTIVITY.1_PORTS_UTIL EXE_ACTIVITY.2_PORTS_UTIL
off-core data reads store-bound cycles backend-bound term backend-bound term backend-bound term
marker, the library reads the counters of a single event set; at the exit marker, it reads them again and accumulates the difference. Torch-PIM collects the following information during profiling. The first is the elapsed time that a region accumulates over all of its invocations, taken from a timing run that executes the nests with the counters disabled. The other two come from the counter run, whose events are listed in Table 1. Off-core data reads give the traffic the region issues past the last-level cache. The remaining four are the counters that the store-bound of the Top-down Microarchitecture Analysis (TMA) [7] is computed from. No counter is multiplexed and no record is scaled up from a sampled fraction of the execution. 4.4
Offload Decision Rule
Torch-PIM decides placement in two stages. The first excludes regions that cannot recover the minimum offloading costs, or whose iterations cannot be distributed across the PIM cores; the second selects those that are memory-bound on the host. offload(𝑟 ) = g_work(𝑟 ) ∧ g_mem(𝑟 )
(1)
Stage 1: Minimum Work. A region 𝑟 passes the first stage only if it satisfies both of the following: g_work(𝑟 ) = (C1) ∧ (C2) (C1) 𝑇cpu (𝑟 ) (C2)
min ≥ 𝑇ctx + 𝑇pim (𝑟 )
𝑁 iter (𝑟 ) ≥ 𝑁 core
(2) (3) (4)
A region that fails either condition is placed on the host without evaluating the second stage. For offloading to pay off, the host execution time must exceed the sum of two expenditures: the cost of switching execution to the memory side, 𝑇ctx ; the cost of moving the operands; and the time the kernel takes on the PIM cores, 𝑇pim (𝑟 ). C1 charges the first and the third. The third cannot be known without executing the region, but it can be bounded from below. A loop body that executes on PIM contains at least one instruction, so 𝑁 total (𝑟 ) iterations issue at least 𝑁 total (𝑟 ) · 𝐼 min instructions; we take 𝐼 min = 1. The target PIM core is singleissue and in-order, so at most one instruction retires per cycle, and distributing the iterations perfectly across 𝑁 core cores
leaves an execution time of at least 𝑁 total (𝑟 )·𝐼 min /(𝑁 core ·𝑓pim ). The quantity sets memory access latency, bank conflicts, and load imbalance to zero, so no execution can fall below it. The threshold in C1 is therefore a lower bound on the true threshold, and a region that fails C1 cannot satisfy the true condition either. Stage 2: Memory Boundedness. The second stage confirms memory boundedness of each loop from the counter run of Section 4.3 g_mem(𝑟 ) = (C3) ∨ (C4) 𝑛 rd (𝑟 ) · 𝐿 > 𝜃 𝑇cpu (𝑟 ) 𝑐 store (𝑟 ) StoreBound(𝑟 ) = > 𝜎 𝑐 backend (𝑟 )
(5)
(C3) 𝐵𝑊read (𝑟 ) =
(6)
(C4)
(7)
Here 𝑛 rd is the number of off-core data read requests, 𝐿 = 64 B is the cache line size, and 𝑐 store is the cycles bound on stores; the backend term is the sum of the three backend-bound events and 𝑐 store , following the definition in the Top-down Microarchitecture Analysis (TMA) [7]. The native events are those in Table 1. The two paths are separate because reads and writes appear in different counters. Read traffic is recorded directly as off-core read requests. These are the requests that leave the last-level cache, so multiplying by the cache line size and dividing by the elapsed time gives the DRAM read bandwidth the region actually achieved. Non-temporal stores bypass the cache and so never appear in the off-core read counter; a region dominated by writes shows 𝐵𝑊read ≈ 0 and cannot pass (6). The write side is therefore detected not by traffic volume but by the stall it produces, namely the fraction of backend cycles bound on stores. Each path covers regions the other cannot observe, and using only one of them leaves the other side entirely undetected. We set the two thresholds as follows. We sweep 𝜃 over 0.50–4.00 GB/s on the 44 profiled loops, of which 29 benefit from offloading under end-to-end CPU/PIM execution. Below 1.73 GB/s, compute-bound regions pass, and false positives rise to 3–4; at and above 1.80 GB/s they reach their minimum of 2. In the other direction, false negatives grow from 6 to 15 at 2.25 GB/s and accuracy falls to 0.614. We therefore adopt 𝜃 = 2.0 GB/s, the midpoint of the [1.80, 2.13] GB/s interval over which false positives are minimal, which is 0.12 × 𝐵 cpu for the 𝐵 cpu = 16.754 GB/s of our platform and is carried to other platforms as that ratio. At this setting, 26 regions are selected, and the decision matches end-to-end execution for 37/44 = 0.841 of them. For 𝜎 we take 0.4, the most conservative value of the range at which TMA classifies a region as store bound [7].
Lee et al.
Table 2. System configuration. Out of order CPU Processor Cores Clock Caches
Intel Xeon Gold 6330 (Ice Lake-SP) 2 sockets × 28 cores, 2 threads/core 2.00 GHz base, 3.10 GHz max 48 kB L1D, 32 kB L1I, 1.25 MB L2 per core; 42 MB L3 per socket
General-purpose in-order PIM cores Cores Clock Caches
32, 64, 128 1 GHz, single-issue 32 kB L1I, 32 kB L1D
The three offloading configurations differ only in the unit at which the decision is made.
Table 3. Benchmark categories in our evaluation. Category labels (a)–(f) correspond to the panels of Figure 3. Category
Benchmarks
Explanation
Tensor operator ew_add, copying, scale, zeroing, relu
Streaming kernels with no reuse; the case offloading should always take.
Transcendental silu, gelu, tanh
Elementwise shape but compute-bound bodies; the case offloading must reject.
Reduction / normalization
softmax, layernorm, rmsnorm
Operators that lower to loop nests of mixed boundedness, requiring per-nest decisions.
GEMM / conv primitive
mac, gemv, matmul, depthwise conv
Reuse-bearing kernels spanning the decision boundary as 𝑁 core varies.
Elementwise / streaming
Layer of tensor operators Composite layer
attention, MLP
Multi-operator layers used to evaluate selectivity beyond single operators.
Real-world workload Real LLM decode layer
GPT-J-6B, LLaMA-7B
5
Evaluation
5.1
Experimental Setup
• CPU-only execution (baseline): no loop is offloaded, and the entire workload remains CPU-resident. • Basic-block-level offloading: the same decision rule is applied to basic blocks, so each block is judged independently. • Operator-level offloading: the same decision rule is applied at the granularity of the operators the model is written in, so that all loop nests in a single operator are governed by the same decision. • Torch-PIM: Our loop-level offloading method.
Decode-layer inference
Benchmarks and comparisons. We evaluated TorchPIM on 15 tensor operators and two layers in PyTorch [31] and two real-world DL models: GPT-J [37] and LLaMA [36]. We compare our method with three different offloading configurations, as explained below.
System. We target PIM architectures that place generalpurpose programmable cores in memory. Table 2 summarizes the heterogeneous system configuration used in the simulation. The CPU side is a real Ice Lake machine, a highperformance server processor, and all host-side times are measured on it bare-metal. The PIM performance is measured on the PIMProf [39] simulator, which is on the Sniper simulator [4]. Each side of the heterogeneous system is measured by the means that is most reliable for it. Regions that remain on the host are measured on real hardware, where cache behavior, prefetching, and memory-level parallelism are those of a real machine rather than a model of one. Regions placed on the PIM have no corresponding hardware and are therefore measured under simulation. We sweep the number of cores in PIM from 32 to 128 in evaluating workloads to identify the impact of parallelism in PIM on performance improvement. Data-movement cost. Torch-PIM assembles the data-movement cost of a placement from what the lowering pipeline collected. Offloading a loop nest to PIM costs three things: the transitions between host and PIM, the bytes of the nest’s ins that must cross to PIM before it runs there, and the bytes of its outs that must cross back afterward. Writing 𝑁 region for the number of host–PIM transitions the placement induces, fetch and flush for those two sets of bytes, and 𝐵 rd and 𝐵 wb for the effective bandwidths at which they move, that cost is 𝑇copy = fetch/𝐵 rd + flush/𝐵 wb .
(8)
The per-allocation sizes and trip counts of Section 4.2 give fetch and flush; the remaining terms do not come from the IR. 𝑁 region is an output of the decision rule, since it depends on which regions the rule places on PIM and where the boundaries between them fall. The PIM cores share DRAM with the host, so the movement 𝐵 rd and 𝐵 wb describe traffic between the host cache hierarchy and PIM. Bank-level aggregate bandwidth on the PIM side is reported at the terabyte-per-second scale on real systems [23], more than an order of magnitude above host channel bandwidth. We therefore use the bandwidth of the host memory system for both, fixed once per target platform.
Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch ew_add copying
scale zeroing
relu silu
10
gelu
tanh
softmax
1.35
1.06
layernorm
rmsnorm
1.30 1.04
8
1.25 1.20
1.02
6
1.15
1.00
1.10
4 0.98
1.05
2
0.96
1.00
0
0.94
0.95
Speedup
(a) Elementwise / streaming
mac
matmul
(b) Transcendental
(c) Reduction / normalization
depthwise conv
gemv
attention
MLP
GPT-J-6B
4.5
1.8
LLaMA-7B
5
4.0 3.5
1.6
4
3.0 1.4
3
2.5 2.0
1.2
2
1.5 1.0
1.0
1
0.5 32
64
128
32
(d) GEMM / conv primitive
64
128
(e) Layer of tensor operators Tensor operator
Layer of tensor operators
32
64
128
(f) Real-world workload
Real-world workload
Figure 3. Speedup across core counts for the kernel benchmarks. This is an upper bound. Equation 8 does not estimate the movement cost but bounds it from above, for two reasons. First, it charges the entire read set of every region placed on PIM to fetch and its entire write set to flush. It does not account for data that stays resident on the PIM side across successive PIM regions, so a buffer whose producer and consumer are both on PIM — and which is therefore never moved at all — is still charged every time. Second, it treats every execution of a region as moving the full amount again: an input that remains unchanged while an enclosing loop iterates is multiplied by the trip count all the same. 5.2
Performance Analysis
In the evaluation, the execution time of a workload, denoted as hybrid execution time 𝑇ℎ𝑦𝑏𝑟𝑖𝑑 , is computed by the following equation, 𝑇ℎ𝑦𝑏𝑟𝑖𝑑 (𝑁 ) = 𝑇ℎ𝑜𝑠𝑡 + 𝑇𝑝𝑖𝑚 (𝑁 ) + 𝑇𝑐𝑜𝑝𝑦 + 𝑁 region𝑇sw,
(9)
where 𝑇ℎ𝑜𝑠𝑡 is the time spent in regions the rule left on the CPU, 𝑇𝑝𝑖𝑚 (𝑁 ) is the execution time of the offloaded regions on 𝑁 PIM cores, and 𝑇𝑐𝑜𝑝𝑦 is the cost of moving their operands between host and PIM memory explained in Section 5.1. For 𝑇sw , we use the 2 𝜇𝑠 adopted by PIMProf [39].
Elementwise/streaming. Every loop nest in this group is offloaded, leaving no host time outside the offloaded regions. Speedup rises from 6.3× at 32 cores to 7.3 × at 128, a 30 % gain for an eightfold increase in cores, and the five kernels of Fig. 3(a) follow the same trajectory. The execution time is the sum of 𝑇𝑝𝑖𝑚 and 𝑇𝑐𝑜𝑝𝑦 as Torch-PIM offloaded the entire loop nests to PIM, and only 𝑇𝑝𝑖𝑚 decreases with the core count. Operand movement accounts for 7.1 % of all-CPU time and is independent of the core count, so each doubling of cores yields a smaller gain than the last. The flattening of these curves comes from the cost of host–PIM data movement, not from a limit on PIM throughput. Transcendental. Fig. 3(b) shows no improvements for workloads in Transcendental. No loop nest in this group is offloaded to PIM. They passed both C1 and C2, because the loop nests perform enough arithmetic per element. However, their measured read bandwidth stays below 𝜃 ; they issue no non-temporal stores, so they satisfy neither C3 nor C4. The whole benchmark therefore runs on the host, and no speedup is observed. Reduction/Normalization. Fig. 3(c) depicts the performance of workloads in Reduction/normalization. softmax
Lee et al.
gains 1.5 % at 128 cores. One of its eight nests is offloaded; the remaining seven, including the nest that evaluates the exponential, stay on the host and account for 97 % of all-CPU time. layernorm and rmsnorm capture 48.5 % and 63.5 % of allCPU time respectively, but behave differently with the core count. rmsnorm improves monotonically, from 0.96 × at 32 cores to 1.30 × at 128, a gain of 35 %. At 32 cores, the operand movement cost 𝑇𝑐𝑜𝑝𝑦 outweighs the PIM gain and the offloaded version runs 4 % slower than the host. layernorm instead peaks at 64 cores, at 1.19 ×, 1.27 ×, and 1.21 × for 32, 64 and 128 cores — a 6.7 % gain followed by a 4.7 % regression. The regression originates on the PIM side: the PIM execution time of the offloaded regions rises by 17.8 % from 64 to 128 cores. The difference lies in how much parallel work the offloaded regions contain. Adding cores shrinks the per-core work but adds synchronization and result-aggregation overhead; the fewer parallel iterations a region has, the smaller the core count at which the overhead wins. layernorm’s offloaded regions are smaller than rmsnorm’s, placing that crossover near 64 cores, whereas rmsnorm does not reach it within 128. GEMM/Convolution Primitives. Fig. 3(d) depicts speedups for the GEMM/convolution primitives. All four benchmarks in this group perform multiply-accumulate, yet the improvement with offloading differs. What separates them is not the operation but the reuse per element. matmul is unchanged at 1.00 ×. An 𝑁 × 𝑁 square GEMM performs 2𝑁 3 FLOPs over 3𝑁 2 four-byte elements, an arithmetic intensity of 𝑁 /6 FLOP/byte, which is 341 at 𝑁 = 2,048; reuse grows with 𝑁 , so larger instances are further from the memory-bound regime. The nest is the largest in this group and passes Stage 1, but its reuse keeps tiles resident in cache, reducing off-core read requests, and its measured bandwidth stays below 𝜃 . The rule thus establishes that the nest is large and parallel enough, then rejects it at Stage 2 for not being memory-bound. gemv performs the same multiply-accumulate at the opposite extreme. Each matrix element is read once, and only the vector is reused, so memory traffic per element is high. Its measured bandwidth exceeds 𝜃 , and the condition C3 holds, improving by 52% from 32 to 128 cores to reach 1.56 ×. The same operation receives opposite verdicts on reuse alone, and the rule arrives at the distinction by observing bandwidth rather than computing reuse statically. In mac, the nest that passes is not the multiply-accumulate loop but the one preceding it. That loop fills a result array sequentially with no reuse and uses non-temporal stores, which bypass the cache, so its write traffic never appears in the off-core read counter. Its measured read bandwidth is near zero, and it cannot pass C3. However, the fraction of backend cycles bound on stores exceeds 𝜎, and it enters
through C4. mac is the only benchmark that passes through C4 alone, and without the write-side path this nest would appear in no counter at all. The benchmark improves by 61 % from 32 to 128 cores, reaching 1.82 ×. For depthwise_conv the speedup is flat at 1.40× across 32, 64 and 128 cores, even though the PIM execution time of the offloaded regions falls in inverse proportion to the core count, halving over the same range. Of the three components of hybrid execution time, only the PIM term varies with the core count, and it already accounts for just 0.7% of all-CPU time. Halving it therefore shortens end-to-end time by 0.35%, which moves the speedup from 1.40× to roughly 1.405× — below the second decimal place. What dominates instead is the host time left outside the offloaded regions: two of the three loop nests are offloaded, capturing 29% of all-CPU time and leaving 71% on the host. This residual is independent of the core count and bounds the attainable speedup at 1.408×; the measured 1.40× reaches 99.4% of that bound. The curve is flat not because PIM stops improving, but because only 0.7% of end-to-end time remains for any such improvement to act on. The residual host time is the bulk convolution loop. Its reuse rate is high enough that the measured read bandwidth stays below 𝜃 , so it does not pass C3 , and it issues no nontemporal stores, so it does not enter through C4 either. This case is therefore a property of the workload rather than a failure of the rule. Every memory-bound loop nest available for offloading was captured, and the result attains 99.4% of what that capture permits, so no further gain is available on the decision side. Layer of Tensor Operators and Real-World Workloads. Fig. 3(e) and 3(f) show the improvements for Layer of tensor operators and Real-world workloads. GPT-J-6B and LLaMA-7B offload 22 % and 24 % of their loop nests, capturing 97.0 % and 96.0 % of all-CPU time. Speedup rises by 45 % from 32 to 128 cores for GPT-J-6B, reaching 5.1 ×, while LLaMA-7B reaches 3.6 ×. Every offloaded GEMM is a thin GEMM, that is, a GEMV. In an 𝑀 × 𝑁 × 𝐾 product with small 𝑀, the weight matrix [𝑁 × 𝐾] is read once and reuse exists only along 𝑀; the arithmetic intensity is 𝑀/2 FLOP/byte, or 0.5 at 𝑀 = 1, which is memory-bound. This follows from the workload rather than from the rule: LLM decoding generates one token at a time, so the batch dimension is thin, and every projection becomes a weight-streaming operation. mlp is the extreme case. At batch size one, every linear layer takes the form 𝑊 · 𝑥; its two offloaded nests have a low reuse rate and measured bandwidths of 6.97 and 9.19 GB/s, three to five times 𝜃 . The bandwidth is high because there is no reuse to exploit. The offloaded sets of GPT-J-6B and LLaMA-7B show the same structure in their measured signatures. One group has bandwidths of 15–18 GB/s with a store-bound fraction of zero — nests 401, 901, 1201, 4001, 4901 and 5801 in GPT-J-6B
Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch
Table 4. Offloaded loop nests per operator as the problem size 𝑁 varies. Each entry is 𝑁 →𝑘, where 𝑘 of the operator’s loop nests are selected for offloading. Operator
Nests
Offloaded nests at each 𝑁
Monotone mac mlp attention gemv relu ew_add
3 7 14 2 1 1
256→0, 1024→0, 2048→1, 4096→2 1024→0, 4096→2, 8192→4 512/1024/2048→0, 4096→3, 8192→3 512/1024→0, 4096/8192/16384→1 512/1024→0, ≥4096→1 512/1024→0, ≥4096→1
Non-monotone softmax
8
1024→0, 4096→2, 32768→4, 65536→3
Size-invariant depthwise_conv
3
2 at all sizes
operator − loop
16384→5,
and 301, 1701, 4501, 5401 and 6301 in LLaMA-7B — corresponding to the 𝑊𝑞 , 𝑊𝑘 , 𝑊𝑣 , 𝑊𝑜 and FFN projections, each streaming a [𝑑 × 𝑑] weight matrix once. The other has bandwidths of 3–9 GB/s with store-bound fractions of 0.07–0.15 and mixes elementwise computation with stores. The 50 and 51 loop nests in GPT-J-6B and LLaMA-7B rejected to be offloaded, and 94% of the rejections occur at Stage 1: most nests are eliminated on size before memory boundedness is evaluated at all. Host time outside the offloaded regions is 3.0 % and 4.0 %, so the remaining gap lies in PIM execution and data movement rather than in regions the rule failed to identify. 5.3
Table 5. Cost of each granularity relative to loop granularity. All entries are percentages of the end-to-end execution time under loop granularity.
Offload results can be different according to input sizes
Table 4 reports the number of loop nests selected for offloading across 11 operators and 54 loop nests as the problem size 𝑁 varies. Two observations follow. First, the decision is made per loop nest, not per operator. In 7 of the 8 operators with more than one nest, only a subset of the nests is selected even at the largest 𝑁 : 3 of 14 for attention, 3 of 12 for layernorm, 3 of 8 for softmax, 4 of 7 for mlp, 2 of 3 for mac, 2 of 3 for depthwise convolution, and 1 of 2 for gemv. A framework that treats an entire operator as a single candidate cannot express this split. Second, the selected nests depend on 𝑁 . For mac, mlp, attention, and gemv, no nest is selected at the smallest sizes, and nests cross the threshold one at a time as 𝑁 grows. For layernorm, depthwise convolution, and silu, the count is essentially fixed over the range measured. softmax is not monotone. The count rises from 0 at 𝑁 = 1024 to 2 at 4096 and 5 at 16384, then falls to 4 at 32768 and 3 at 65536. Its 8 nests flip at different values of 𝑁 , so a larger problem size does not uniformly favor offloading.
Benchmark GPT-J LLaMA attention
5.4
BB − loop
Δ𝑇pim (%)
Total (%)
Δ𝑇copy (%)
Total (%)
+9.3 +6.2 +77.9
+9.8 +6.5 +67.8
+4.3 +3.2 +1.6
+4.8 +3.5 +2.7
Decomposing the Performance Gap Across Granularities
Table 5 compares the three granularities. The basic block (BB) granularity and operator-level granularity lose for different reasons: operator granularity raises 𝑇pim from 29.5% to 38.8% of the loop-granularity baseline, while BB granularity leaves 𝑇pim and 𝑇cpu at their loop-granularity values and raises 𝑇copy from 2.1% to 6.4%. Operator granularity places every loop nest of an operator on PIM as a unit, so compute-bound loop nests inside an operator are offloaded along with the memory-bound ones. On GPT-J this gives Δ𝑇pim = 9.3 %, Δ𝑇cpu = −1.1 % and Δ𝑇copy = 1.6 %, summing to 9.8 %, of which 95 % is Δ𝑇pim : loop nests that consume 1.1 % of the baseline on the CPU consume 9.3 % on PIM, a factor of 8.8 ×. Data movement accounts for 16 %, so the cost comes from the placement decision rather than from moving data. On attention the same effect is 77.9 %, of which 10.1 percentage points are recovered on the CPU and copy sides, leaving 67.8 %. In the case of BB granularity, crossings rise from 18 to 372, giving Δ𝑇copy = +4.3% and Δ𝑇switch = +0.4% over 354 additional crossings at 2 𝜇s each, for +4.8%; 91% of this is boundary tensor materialization and 9% the switching constant on GPT-J. The per-crossing copy cost is 21.6 𝜇s on GPT-J against 3.2 𝜇s on attention, scaling with boundary tensor size, so the fragmentation cost is the product of crossing count and boundary tensor size, and it is this product that segment merging keeps bounded.
6
Limitations and Future Work
Empirical basis of the thresholds. 𝜃 and 𝜎 are the only calibrated constants in the rule, and both rest on a narrow base. We sweep 𝜃 on the same 44 loop nests over which we report the 37/44 = 0.841 agreement, and hold out no separate set, so that figure states how well a single threshold separates this population rather than how the calibrated rule transfers to nests it was not fitted on. Two false positives persist at every 𝜃 ≥ 1.80 GB/s and do not disappear up to 4.00 GB/s, which indicates that measured read bandwidth alone does not separate every region. 𝜎 rests on a single observation: mac is the only benchmark in which a nest issues non-temporal stores and therefore the only one that
Lee et al.
enters through C4, so the write-side path is calibrated on one sample. Establishing the portability of the ratio 𝜃 /𝐵 cpu across hosts with different memory systems, and combining the two signals of C3 and C4 into a single decision rather than a disjunction of two independently calibrated thresholds, are both left to future work. Direction of the movement cost. Equation 8 bounds the movement cost from above rather than estimating it. A buffer that a producer writes and a consumer reads while both remain on PIM is charged at every region boundary, and an operand that does not change while an enclosing loop iterates is charged once per trip. The error is one-sided, so the speedups we report are lower bounds on what the same placement would achieve under an exact accounting. Narrowing the bound requires tracking residency across consecutive PIM regions and identifying loop-invariant operands, neither of which the current pipeline does. Torch-PIM emits offload markers and leaves PIM-side code generation to the downstream toolchain, so the PIM execution times we report reflect the quality of that toolchain as well as the placement decision. Scope of the candidate space. Torch-PIM requires every tensor extent to be a compile-time constant, which holds for models exported at a fixed input shape but excludes models whose shapes are resolved at run time. Within such a model, candidates are the loops that bufferization materializes as scf.parallel. A loop that is sequential as emitted but would become parallel under a transformation never enters the candidate space, since the pipeline decides placement over the loops it produces rather than over the loops it could produce. Broadening the candidate space would require legality analysis for the transformations concerned and a decision rule that ranks candidates the pipeline has not yet materialized.
7
Conclusion
Torch-PIM lowers a PyTorch model through MLIR and decides host-versus-PIM placement over the loop nests that progressive lowering materializes. Nothing designates these candidates in advance: they are neither a list of operator types, nor patterns identified beforehand, nor regions a programmer annotated. Each candidate is judged in two stages, on the amount of work it carries and on its memory boundedness, and every quantity the rule consumes is measured on the host, so no candidate is ever executed or simulated on the PIM device. Our evaluations on PIM configurations of 32 to 128 cores reach speedups of up to 8.6 × on tensor operators, 2.9 × on MLP, 4.4 × on Attention, 5.1 × on GPT-J6B, and 3.6 × on decoder layers of LLaMA-7B over CPU-only execution.
References [1] Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi. 2015. PIM-enabled instructions: a low-overhead, locality-aware processingin-memory architecture. In Proceedings of the 42nd Annual International Symposium on Computer Architecture (Portland, Oregon) (ISCA ’15). Association for Computing Machinery, New York, NY, USA, 336–348. doi:10.1145/2749469.2750385 [2] Michael Anderson, Benny Chen, Stephen Chen, Summer Deng, Jordan Fix, Michael Gschwind, Aravind Kalaiah, Changkyu Kim, Jaewon Lee, Jason Liang, Haixin Liu, Yinghai Lu, Jack Montgomery, Arun Moorthy, Satish Nadathur, Sam Naghshineh, Avinash Nayak, Jongsoo Park, Chris Petersen, Martin Schatz, Narayanan Sundaram, Bangsheng Tang, Peter Tang, Amy Yang, Jiecao Yu, Hector Yuen, Ying Zhang, Aravind Anbudurai, Vandana Balan, Harsha Bojja, Joe Boyd, Matthew Breitbach, Claudio Caldato, Anna Calvo, Garret Catron, Sneh Chandwani, Panos Christeas, Brad Cottel, Brian Coutinho, Arun Dalli, Abhishek Dhanotia, Oniel Duncan, Roman Dzhabarov, Simon Elmir, Chunli Fu, Wenyin Fu, Michael Fulthorp, Adi Gangidi, Nick Gibson, Sean Gordon, Beatriz Padilla Hernandez, Daniel Ho, Yu-Cheng Huang, Olof Johansson, Shishir Juluri, Shobhit Kanaujia, Manali Kesarkar, Jonathan Killinger, Ben Kim, Rohan Kulkarni, Meghan Lele, Huayu Li, Huamin Li, Yueming Li, Cynthia Liu, Jerry Liu, Bert Maher, Chandra Mallipedi, Seema Mangla, Kiran Kumar Matam, Jubin Mehta, Shobhit Mehta, Christopher Mitchell, Bharath Muthiah, Nitin Nagarkatte, Ashwin Narasimha, Bernard Nguyen, Thiara Ortiz, Soumya Padmanabha, Deng Pan, Ashwin Poojary, Ye, Qi, Olivier Raginel, Dwarak Rajagopal, Tristan Rice, Craig Ross, Nadav Rotem, Scott Russ, Kushal Shah, Baohua Shan, Hao Shen, Pavan Shetty, Krish Skandakumaran, Kutta Srinivasan, Roshan Sumbaly, Michael Tauberg, Mor Tzur, Sidharth Verma, Hao Wang, Man Wang, Ben Wei, Alex Xia, Chenyu Xu, Martin Yang, Kai Zhang, Ruoxi Zhang, Ming Zhao, Whitney Zhao, Rui Zhu, Ajit Mathews, Lin Qiao, Misha Smelyanskiy, Bill Jia, and Vijay Rao. 2021. First-Generation Inference Accelerator Deployment at Facebook. arXiv:2107.04140 [cs.AR] doi:10.48550/arXiv.2107.04140 [3] Amirali Boroumand, Saugata Ghose, Berkin Akin, Ravi Narayanaswami, Geraldo F. Oliveira, Xiaoyu Ma, Eric Shiu, and Onur Mutlu. 2021. Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks. In 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT). IEEE, 159–172. doi:10.1109/pact52795.2021.00019 [4] Trevor E. Carlson, Wim Heirman, Stijn Eyerman, Ibrahim Hur, and Lieven Eeckhout. 2014. An Evaluation of High-Level Mechanistic Core Models. ACM Transactions on Architecture and Code Optimization (TACO), Article 5, 23 pages. doi:10.1145/2629677 [5] Dan Chen, Hai Jin, Long Zheng, Yu Huang, Pengcheng Yao, Chuangyi Gui, Qinggang Wang, Haifeng Liu, Haiheng He, Xiaofei Liao, and Ran Zheng. 2022. A General Offloading Approach for Near-DRAM Processing-In-Memory Architectures. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 246–257. doi:10. 1109/IPDPS53621.2022.00032 [6] Jinfan Chen, Juan Gómez-Luna, Izzat El Hajj, Yuxin Guo, and Onur Mutlu. 2023. SimplePIM: A Software Framework for Productive and Efficient Processing-in-Memory. In Proceedings of the 32nd International Conference on Parallel Architectures and Compilation Techniques (Vienna, AE, Austria) (PACT ’23). IEEE Press, 99–111. doi:10.1109/ PACT58117.2023.00017 [7] Intel Corporation. 2025. Intel VTune Profiler Cookbook, Top-down Microarchitecture Analysis Method (2025.4 ed.). Retrieved Sept. 10, 2026 from https://www.intel.com/content/www/us/en/docs/vtuneprofiler/cookbook/2025-4/top-down-microarchitecture-analysismethod.html
Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch [8] Nika Mansouri Ghiasi, Nandita Vijaykumar, Geraldo F. Oliveira, Lois Orosa, Ivan Fernandez, Mohammad Sadrosadati, Konstantinos Kanellopoulos, Nastaran Hajinazar, Juan Gomez Luna, and Onur Mutlu. 2023. ALP: Alleviating CPU-Memory Data Movement Overheads in Memory-Centric Systems . IEEE Transactions on Emerging Topics in Computing 11, 02 (April 2023), 388–403. doi:10.1109/TETC.2022. 3226132 [9] Amir Gholami et al. 2024. AI and Memory Wall. IEEE Micro 44, 3 (2024), 33–39. doi:10.1109/MM.2024.3373763 [10] Juan Gómez-Luna, Yuxin Guo, Sylvan Brocard, Julien Legriel, Remy Cimadomo, Geraldo F. Oliveira, Gagandeep Singh, and Onur Mutlu. 2023. Evaluating Machine Learning Workloads on Memory-Centric Computing Systems. In 2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 35–49. doi:10.1109/ispass57527.2023.00013 [11] Ramyad Hadidi, Lifeng Nai, Hyojong Kim, and Hyesoon Kim. 2017. CAIRO: A Compiler-Assisted Technique for Enabling InstructionLevel Offloading of Processing-In-Memory. ACM Trans. Archit. Code Optim. 14, 4, Article 48 (Dec. 2017), 25 pages. doi:10.1145/3155287 [12] Kevin Hsieh, Eiman Ebrahimi, Gwangsun Kim, Niladrish Chatterjee, Mike O’Connor, Nandita Vijaykumar, Onur Mutlu, and Stephen W. Keckler. 2016. Transparent offloading and mapping (TOM): enabling programmer-transparent near-data processing in GPU systems. SIGARCH Comput. Archit. News 44, 3 (June 2016), 204–216. doi:10.1145/3007787.3001159 [13] Karam Hwang. 2026. Samsung LPDDR5X-PIM: World’s First LPDDR based Processing in Memory (PIM) Solution for AI Inference. In 2026 IEEE Hot Chips 38 Symposium (HCS). Samsung Electronics, Stanford, CA, USA. https://hc2026.hotchips.org/program/conference/ Accessed: 2026-09-03. [14] Heike Jagode, Anthony Danalis, Giuseppe Congiu, Daniel Barry, Anthony Castaldo, and Jack J. Dongarra. 2025. Advancements of PAPI for the exascale generation. The International Journal of High Performance Computing Applications 39, 2 (2025), 251–268. doi:10.1177/ 10943420241303884 [15] JEDEC Solid State Technology Association. 2025. LPDDR6 Standard. Technical Report JESD209-6. JEDEC Solid State Technology Association. https://www.jedec.org/standards-documents/docs/jesd209-6 JEDEC [16] JEDEC Solid State Technology Association. 2026. Previews LPDDR6 Roadmap Expanding LPDDR into Data Centers and Processing-in-Memory. Press release. https: //www.jedec.org/news/pressreleases/jedec-previews-lpddr6roadmap-expanding-lpddr-data-centers-and-processing-memory Accessed: 2026-09-03. [17] Qingcai Jiang, Shaojie Tan, Junshi Chen, and Hong An. 2024. A3PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader. In Proc. DATE. 1–6. doi:10.23919/DATE58400.2024.10546698 [18] Rupinder Kaur, Arghavan Asad, and Farah Mohammadi. 2024. A Comprehensive Review of Processing-in-Memory Architectures for Deep Neural Networks. Computers 13, 7 (2024). doi:10.3390/ computers13070174 [19] Asif Ali Khan, Hamid Farzaneh, Karl Friedrich Alexander Friebel, Clément Fournier, Lorenzo Chelini, and Jeronimo Castrillon. 2025. CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4 (Hilton La Jolla Torrey Pines, La Jolla, CA, USA) (ASPLOS ’24). Association for Computing Machinery, New York, NY, USA, 31–46. doi:10.1145/3622781.3674189 [20] Kamil Khan, Sudeep Pasricha, and Ryan Gary Kim. 2020. A Survey of Resource Management for Processing-In-Memory and Near-Memory Processing Architectures. Journal of Low Power Electronics and Applications 10, 4 (2020). doi:10.3390/jlpea10040030
[21] Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, Jeongbin Kim, Jaewook Lee, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyungdeok Lee, Seho Kim, Daehan Kwon, Seongju Lee, Kyuyoung Kim, Sanghoon Oh, Joonhong Park, Gimoon Hong, Dongyoon Ka, Kyudong Hwang, Jeongje Park, Kyeongpil Kang, Jungyeon Kim, Junyeol Jeon, Myeongjun Lee, Minyoung Shin, Minhwan Shin, Jaekyung Cha, Changson Jung, Kijoon Chang, Chunseok Jeong, Euicheol Lim, Il Park, Junhyun Chun, and Sk Hynix. 2022. System Architecture and Software Stack for GDDR6-AiM. In 2022 IEEE Hot Chips 34 Symposium (HCS). 1–25. doi:10.1109/HCS55958.2022.9895629 [22] Young-Cheon Kwon et al. 2021. 25.4 A 20nm 6GB Function-In-Memory DRAM, Based on HBM2. In 2021 IEEE International Solid-State Circuits Conference (ISSCC), Vol. 64. 350–352. doi:10.1109/ISSCC42613.2021. 9365862 [23] Zheyuan Lai and Yingming Pu. 2025. Prim: Principle-inspired material discovery through multi-agent collaboration. (2025). arXiv:2504.08810 [cs.LG] doi:10.48550/arXiv.2504.08810 [24] Chuangshi Liu and Xianfeng Li. 2021. LooPIM: A Loop-Oriented Acceleration Framework for Processing-in-Memory. Journal of Physics: Conference Series 1914, 1 (2021). doi:10.1088/1742-6596/1914/1/012023 [25] Jiantao Liu, Minxuan Zhou, Yue Pan, Chien-Yi Yang, Lana Josipović, and Tajana Rosing. 2025. OptiPIM: Optimizing Processing-in-Memory Acceleration Using Integer Linear Programming. In Proc. ISCA (ISCA ’25). Association for Computing Machinery, New York, NY, USA, 867–883. doi:10.1145/3695053.3731041 [26] LLVM. 2026. Torch-MLIR. https://github.com/llvm/torch-mlir Commit d22e77c. Accessed: Sept. 10, 2026. [27] Satanu Maity, Mayank Goel, and Manojit Ghose. 2023. Data Locality Aware Computation Offloading in Near Memory Processing Architecture for Big Data Applications. In 2023 IEEE 30th International Conference on High Performance Computing, Data, and Analytics (HiPC). IEEE, 288–297. doi:10.1109/hipc58850.2023.00019 [28] Satanu Maity, Mayank Goel, and Manojit Ghose. 2025. CoaT: <u>Co</u>mpiler-<u>A</u>ssisted <u>T</u>wo-Stage Offloading Approach for Data-Intensive Applications Under NMP Framework. IEEE Transactions on Emerging Topics in Computing 13, 3 (2025), 753–767. doi:10.1109/tetc.2024.3495218 [29] Lifeng Nai, Ramyad Hadidi, Jaewoong Sim, Hyojong Kim, Pranith Kumar, and Hyesoon Kim. 2017. GraphPIM: Enabling InstructionLevel PIM Offloading in Graph Computing Frameworks. In Proc. HPCA. 457–468. doi:10.1109/HPCA.2017.54 [30] Geraldo F. Oliveira, Juan Gómez-Luna, Lois Orosa, Saugata Ghose, Nandita Vijaykumar, Ivan Fernandez, Mohammad Sadrosadati, and O. Mutlu. 2021. DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks. IEEE Access 9 (2021), 134457–134502. doi:10.1109/access.2021.3110993 [31] Adam Paszke et al. 2019. PyTorch: An Imperative Style, HighPerformance Deep Learning Library. In Proc. NeurIPS. doi:10.5555/ 3454287.3455008 [32] Youngjin Ro et al. 2022. Aquabolt-XL HBM2-PIM, LPDDR5-PIM With In-Memory Processing, and AXDIMM With Acceleration Buffer. IEEE Micro 42, 3 (2022), 20–30. doi:10.1109/MM.2022.3164651 Samsung Research. [33] Yongwon Shin, Dookyung Kang, and Hyojin Sung. 2025. ATiM: Autotuning Tensor Programs for Processing-in-DRAM. In Proceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing Machinery, New York, NY, USA, 899–915. doi:10.1145/3695053.3731096 [34] Yongwon Shin, Juseong Park, Sungjun Cho, and Hyojin Sung. 2023. PIMFlow: Compiler and Runtime Support for CNN Models on Processing-in-Memory DRAM. In Proceedings of the 21st ACM/IEEE
Lee et al. International Symposium on Code Generation and Optimization (Montréal, QC, Canada) (CGO ’23). Association for Computing Machinery, New York, NY, USA, 249–262. doi:10.1145/3579990.3580009 [35] Xiaotian Sun, Xinyu Wang, Wanqian Li, Yinhe Han, and Xiaoming Chen. 2025. PIMCOMP: An End-to-End DNN Compiler for ProcessingIn-Memory Accelerators. Trans. Comp.-Aided Des. Integ. Cir. Sys. (2025), 1745–1759. doi:10.1109/TCAD.2024.3496847 [36] Hugo Touvron et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288 [cs.CL] doi:10.48550/arXiv.2307.09288 [37] Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/ mesh-transformer-jax. [38] Haonan Wang, Xuxin Xiao, Mingyu Yan, Zhuoyuan Zhu, Dengke Han, Duo Wang, Wenming Li, Xiaochun Ye, Cunchen Hu, Hongyang Chen, and Guangyu Sun. 2025. A Systematic Characterization of LLM Inference on GPUs. arXiv:2512.01644 [cs.AR] doi:10.48550/arXiv.2512. 01644 [39] Yizhou Wei, Minxuan Zhou, Sihang Liu, Korakit Seemakhupt, Tajana Rosing, and Samira Khan. 2022. PIMProf: an automated program profiler for processing-in-memory offloading decisions. In Proceedings of the 2022 Conference & Exhibition on Design, Automation & Test in Europe (Antwerp, Belgium) (DATE ’22). European Design and Automation Association, Leuven, BEL, 855–860. doi:10.23919/DATE54114. 2022.9774560
[40] Sheng Xu, Chun Li, Le Luo, Wu Zhou, Liang Yan, and Xiaoming Chen. 2025. Identifying Optimal Workload Offloading Partitions for CPUPIM Graph Processing Accelerators. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 33, 4 (2025), 1053–1064. doi:10.1109/ TVLSI.2025.3526201 [41] Liang Yan, Mingzhe Zhang, Rujia Wang, Xiaoming Chen, Xingqi Zou, Xiaoyang Lu, Yinhe Han, and Xian-He Sun. 2021. CoPIM: A Concurrency-aware PIM Workload Offloading Architecture for Graph Applications. In Proc. ISLPED. 1–6. doi:10.1109/ISLPED52811.2021. 9502483 [42] Peiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati, Onur Mutlu, Gennady Pekhimenko, and Christina Giannoula. 2026. DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures. In 2026 ACM/IEEE 53rd Annual International Symposium on Computer Architecture (ISCA). IEEE, 2582–2599. doi:10.1109/ISCA66397.2026.00180 [43] Zhisheng Ye, Wei Gao, Qinghao Hu, Peng Sun, Xiaolin Wang, Yingwei Luo, Tianwei Zhang, and Yonggang Wen. 2024. Deep Learning Workload Scheduling in GPU Datacenters: A Survey. Comput. Surveys 56, 6 (Jan. 2024), 1–38. doi:10.1145/3638757 [44] Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer. 2024. LLM Inference Unveiled: Survey and Roofline Model Insights. arXiv:2402.16363 [cs.CL] doi:10.48550/arXiv.2402.16363