arXiv:2609.24628v1 [cs.DC] 21 Sep 2026
Bridging the Vendor Gap: Enabling AMD GPU Support for Awkward Array via ROCm/HIP for the HL-LHC Era Ianna Osborne1,∗ , Maxym Naumchyk1 , Tai Sakuma1 , Andres Rios-Tascon1 , and Peter Elmer1 1
Princeton University, Princeton, NJ 08544, USA, on behalf of the Institute for Research and Innovation in Software for High Energy Physics (IRIS-HEP) Abstract. The High-Luminosity LHC (HL-LHC) will demand order-of-
magnitude gains in analysis throughput, and increasingly those gains must come from GPUs that are not made by a single vendor. Leadership-class systems such as El Capitan, Frontier and LUMI are built on AMD accelerators, yet the Scikit-HEP analysis stack—and Awkward Array in particular—has grown up CUDA-first. We report on rawkward, a Rust-backed kernel engine that adds a ROCm/HIP backend for Awkward Array’s nested, jagged, variable-length data structures. Our central finding is that a naïve source-level port of CUDA kernels to HIP loses 5–10× in performance on irregular kernels, because AMD’s 64-lane wavefronts, higher register pressure and more expensive divergence behave fundamentally differently from NVIDIA’s 32-thread warps. We show that a small, reusable set of optimization patterns—loop flattening, 128-bit vectorized loads, splitting fused kernels, and profile-guided launch configuration— recovers CUDA-class performance without changing the public API. A Rust macro-and-match dispatch layer keeps a single, backend-agnostic call site while emitting vendor-specific kernel strategies, and the type system enforces buffersize and lifetime correctness at compile time. On a two-socket AMD Instinct MI210 node we measure GPU speedups from 1.03× (bandwidth-bound sum) up to 12.5× (count) over 128 EPYC 7763 CPU cores, and the Rust CPU kernels match or beat on aggregate the incumbent C++ kernels (geometric-mean runtime ratio 0.37× across twelve kernels). We argue that these patterns constitute a practical recipe for performance-portable, vendor-agnostic HEP analysis kernels.
1 Introduction The physics programme of the High-Luminosity LHC (HL-LHC) rests on a sharp increase in instantaneous luminosity, and with it an order-of-magnitude increase in the volume and complexity of the data that end-user analyses must process [1, 2]. CPU scaling alone will not close that gap; GPU acceleration has emerged as the only viable path to the throughput that HL-LHC analysis will require [3]. The question is no longer whether to use GPUs, but whose GPUs. For most of the last decade the answer, in practice, was NVIDIA. The Scikit-HEP ecosystem, and Awkward Array [4] at its centre, developed a CUDA-first culture: CUDA kernels, CUDA continuous integration, and CUDA benchmarks, with the rest of the tooling following ∗ e-mail: [email protected]
suit. That choice was reasonable when the accelerated systems available to physicists were overwhelmingly NVIDIA-based. It is no longer a safe assumption. Several of the largest systems now available to the field are built on AMD accelerators—El Capitan, Frontier and LUMI among them—and the coming MI300A-class allocations put AMD GPUs directly in front of HEP workflows. A stack that runs only on one vendor’s hardware leaves a growing fraction of the world’s fastest machines unused. This paper reports on our effort to close that gap for Awkward Array. We introduce a ROCm/HIP [8] backend with a Metal backend as a second non-NVIDIA target, developed inside a dedicated Rust engine we call rawkward, that supports the nested, jagged and irregular data layouts that make HEP data distinctive, and we benchmark it on AMD Instinct MI210/MI250 (CDNA2) hardware with an eye toward forward-compatibility with MI350 (CDNA4) and its 288 GB of HBM3E memory. Our contributions are: • a characterization of the “vendor gap” for irregular kernels, showing quantitatively that direct CUDA→HIP translation costs 5–10× on the memory-bound, divergence-sensitive kernels that dominate jagged-array processing (Sect. 3–5); • a small, reusable catalogue of HIP optimization patterns that recover CUDA parity—loop flattening, vectorized 128-bit loads, kernel splitting over fusion, and profile-guided launch configuration (Sect. 5, 6); • a Rust orchestration layer whose macro-based dispatch keeps one backend-agnostic API across CUDA and HIP while the type system enforces memory safety at compile time (Sect. 4); and • a statistics-driven microbenchmark on a two-socket MI210 node, reporting per-kernel GPU-versus-CPU speedups and a Rust-versus-C++ kernel comparison (Sect. 7–8). The goal is not simply to make Awkward Array run on AMD hardware, but to do so at a performance envelope comparable to its CUDA equivalents, and to distil the experience into guidance that generalizes beyond a single library.
2 Background and motivation 2.1 Irregular data and Awkward Array
HEP event data is irregular by design: the number of jets, tracks or leptons varies event by event, records nest inside records, and lists have variable length. Awkward Array [4] is the Pythonic interface for exactly this structure, presenting ragged lists and nested records as first-class, NumPy-like arrays while storing them in a flat, columnar layout. That columnar representation is what makes vectorized and GPU execution possible: the logical structure is carried by offset and index buffers, and the payload lives in contiguous content buffers. Nearly every operation on such data—slicing, masking, reduction over sublists—reduces to manipulating those index buffers, which is where the performance challenges concentrate. 2.2 Why GPUs, and why portability
The luminosity increase at the HL-LHC demands throughput gains that CPU scaling cannot supply on its own [1]. GPU acceleration of end-user analysis has therefore moved from experiment to expectation [2, 3], and related work has begun to accelerate Awkward Array itself and the schemas that feed it, for example through NVIDIA’s CCCL and through coffea schema modifications [5, 6]. The CUDA ecosystem for this work is mature. ROCm/HIP is catching up, but the maturity gap hides a subtler trap: a HIP port that compiles and runs is
Table 1. Architectural mental models. The differences that matter for jagged-array kernels concentrate in wavefront width, the cost of divergence, and register pressure.
Execution width
Divergence
NVIDIA / CUDA
AMD / HIP
warp = 32 threads; more warps in flight hide latency across independent work relatively cheap; only inactive lanes stall on a branch split
wavefront = 64 lanes executing in lockstep
Scattered loads
out-of-order schedulers hide irregular access; naïve gathers often still reach peak bandwidth
Vectorization
helpful
expensive; a branch split can lose half the throughput of a 64-wide wavefront register file shared across more lanes; spills hit LDS/global memory hard, so occupancy tuning is critical critical; double-width SIMD and explicit LDS layout unlock full throughput
not the same as a HIP port that performs. On irregular kernels, naïve ports routinely leave 5– 10× of performance on the table. Performance portability across vendors—one source, parity performance on each target—remains substantially unsolved, and it is the reason rawkward exists.
3 The vendor gap for irregular kernels CUDA kernels for Awkward Array exist and have been optimized. HIP is a different story— not because ROCm is immature as a toolchain, but because the underlying execution model differs in ways that matter precisely for irregular access patterns. Table 1 summarizes the two mental models. The consequence is concrete. A 64-lane wavefront that diverges pays roughly twice the penalty of a 32-thread warp, and the wider register file is shared among more lanes, so the same kernel that fits comfortably on NVIDIA hardware may spill on AMD and collapse occupancy. Scattered-gather patterns that NVIDIA’s memory schedulers tolerate can thrash the AMD cache hierarchy. None of this is visible at the source level: the semantic gap between “compiles” and “performs” is exactly the gap this work measures and closes. Our approach is therefore HIP-specific optimization, not translation—tuning each kernel for wavefront width, LDS layout and occupancy—while keeping the public API unchanged so that call sites never see the difference.
4 Rust as a portable orchestrator Awkward Array’s kernels have historically been generated: the C++ backend generation is highly repetitive, which is manageable for one target but multiplies awkwardly across vendors. We chose instead to express the target-dependent dispatch in Rust, using a macro-andmatch pattern that handles the vendor split cleanly. rawkward is organized as a “foundry”: a clean-room engine that mirrors the existing kernel structure closely enough that validated Rust kernels can be copied back into Awkward Array with minimal refactoring, while giving the development a modern toolchain and clear language boundaries. A guiding principle is
that kernels do not require Python: the Rust core is a standalone compute layer, exposed to Python through a lean PyO3 binding, but independent of the interpreter’s lifecycle. Three properties make Rust a good fit for the orchestration role. First, the same Rust API serves both CUDA and HIP: call sites are backend-agnostic, and the same function signature compiles and dispatches correctly on each target. The dispatch is designed to be vendorparametric; the HIP backend is implemented in Rust here, while the CUDA path is Awkward Array’s existing (CuPy) backend. Second, the backend—not the caller—decides kernel strategy: the dispatch match arms select fused-versus-split and other vendor-specific choices, and the caller never sees that difference. Third, Rust’s ownership, lifetime and type rules catch data races and mismatched buffer sizes at compile time rather than at run time. The dispatch itself is a small macro (Listing 1); every call routes through it, so HIP and CUDA diverge only where the hardware demands it, with zero per-target boilerplate at the call site. Listing 1. Macro-based backend dispatch. A single match routes each call to the correct vendor backend; the caller is backend-agnostic. macro_rules! dispatch { ($backend:expr, $fn:ident, $($arg:expr),*) => { match $backend { Backend::Hip => hip::$fn($($arg),*), Backend::Cuda => cuda::$fn($($arg),*), } } }
Concretely, the HIP arm applies vectorized loads, loop flattening and a split-kernel strategy—each technique aimed directly at the 64-lane wavefront—while the CUDA arm fuses stages for warp-friendly behaviour and relies on the 32-thread scheduler to tolerate divergence. Because the strategy lives behind the dispatch macro, adding a third backend later (Sect. 9) is a matter of adding a match arm, not rewriting kernel call sites.
5 Kernel case studies 5.1 The carry kernel
The carry kernel is the core primitive of jagged indexing. Given a carry array of indices, it gathers from[carry[i]] into to[i]—a pure gather by indirection. Almost every nestedlist operation ultimately reduces to a carry, which makes it the right kernel to study first. It is also the hardest for AMD hardware: it is memory-bound, irregular and divergence-sensitive. The random-access gather defeats cache prefetchers, and adjacent threads touch non-adjacent memory. On CUDA the naïve implementation performs well. The obvious scatter loop compiles and ships; 32-thread warps combined with hardware scatter-gather tolerate the irregular access and keep the memory pipeline full even when adjacent threads reach distant addresses. The same source ported directly to HIP compiles and runs, but runs 5–10× slower: 64-lane wavefronts amplify divergence, and the gather pattern thrashes the L2 cache with as many simultaneous cache-miss streams as there are lanes. The semantic gap is invisible in the source—the two kernels look identical. Recovering parity takes three targeted changes, none of which touches the API: • Loop flattening: hoist conditionals out of the inner loop so that all 64 lanes execute the same instruction sequence, eliminating intra-wavefront divergence stalls. • float4 vectorized loads: pack four 32-bit reads into one 128-bit transaction, cutting transaction count by 4× and saturating HBM bandwidth instead of thrashing the cache.
• Lower register pressure: reorder operands and reduce the number of live variables, raising occupancy so the wavefront scheduler stays busy. With these applied, the HIP carry kernel reaches the same performance envelope as its CUDA equivalent [5]—the same class of speedup. 5.2 Fusion or split
A second, more surprising divergence between the vendors concerns kernel fusion. On CUDA, fusing two stages A + B into a single kernel is typically faster: intermediate values stay in registers or L1 between stages, avoiding a global-memory round-trip, and the 32-thread warp scheduler handles the fused control flow gracefully. On HIP the same choice is often counter-productive. Fusing stages merges their register budgets across 64-lane wavefronts; occupancy drops and stalls multiply. Two clean, branch-free passes—each of which fits the wavefront register file and runs at full occupancy—beat one fused pass. The rule inverts: fuse on CUDA, split on HIP. Encoding that choice in the dispatch backend, rather than in the caller, is precisely what lets one API serve both.
6 Optimization patterns Across the kernel set, the same four patterns recur, and together they form a practical recipe for HIP parity on irregular kernels: 1. Loop flattening (carry, argsort): eliminate intra-wavefront branches so every lane runs the same instruction stream. 2. float4 vectorized loads (carry, reductions): four 32-bit reads become one 128-bit transaction, cutting L2 pressure 4× and saturating HBM2e bandwidth. 3. Split over fused kernels (fusion, argmax, sum): two lean passes at full occupancy beat one register-hungry fused pass on HIP. 4. Rust macro dispatch (all kernels): one match-based macro routes each call to the right vendor backend, so HIP and CUDA diverge only where the hardware demands it. These are deliberately low-level and mechanical, which is the point: they can be applied kernel-by-kernel, validated in isolation, and carried across libraries.
7 Benchmarking methodology Reliable microbenchmarks are notoriously easy to get wrong, so we drive all measurements from a statistics-driven harness in Rust (the Criterion framework [10]), reporting the mean over 100 samples per configuration rather than a single timed run. Each kernel is exercised at three input sizes chosen to stress the L1, L2 and L3/main-memory regimes respectively, so that cache effects and the GPU launch floor are visible rather than averaged away. The GPU measurements include the PCIe copy, so the comparison is end-to-end rather than deviceresident only. Benchmark results and scripts are archived publicly [7]. The primary test node, della-milan, is summarized in Table 2. It pairs two AMD EPYC 7763 (Milan / Zen 3) sockets with two AMD Instinct MI210 (CDNA2, gfx90a [9]) accelerators, so that a fair CPU-versus-GPU comparison can be made on the same host, and so that the GPU path exercises the discrete, non-unified-memory configuration typical of current deployments. A subset of the kernels is additionally cross-checked on Apple M4 Metal hardware, confirming that the engine’s portability is not specific to a single non-NVIDIA target.
Table 2. The della-milan benchmark node. CPU and GPU share the same host, enabling an end-to-end comparison that includes the PCIe transfer. CPU
GPU
Device
2× AMD EPYC 7763 (Milan / Zen 3)
Parallelism
128 cores total (64/socket, no HT), 3.53 GHz boost, AVX2 ∼1 TB DDR4, L2 64 MiB, L3 512 MiB, 8 NUMA nodes glibc, ROCm host
2× AMD Instinct MI210 (CDNA2, gfx90a) 104 CUs, 6 656 shaders
Memory Notes
count countnonzero argmax argmin max min prod sum
64 GB HBM2e, ∼1.6 TB/s PCIe (discrete, no unified memory)
12.5× 9.5× 5.6× 5.4× 5.2× 5.0× 4.8× 1.03×
0
parity
2 4 6 8 10 12 Speedup (CPU time / GPU time), large size
Figure 1. Per-kernel speedup of the HIP backend on 2× MI210 against 128 EPYC 7763 cores at the large input size (higher is better; the dashed line marks parity). Write-dominated kernels win most; the bandwidth-bound sum reaches only parity because AVX2 already saturates CPU memory bandwidth.
8 Results 8.1 GPU versus CPU on irregular kernels
Figure 1 shows per-kernel GPU speedups over the full 128-core CPU at the large input size. The kernels fall into clear tiers. count achieves the best result, 12.5×: it is a write-only pattern with no arithmetic, so the GPU issues a single read per segment and is limited only by memory throughput. countnonzero, a predicated write, follows at 9.5×. The argumentand value-reductions—argmax, argmin, max, min and prod—are compare-reduce, memorybound kernels and land in the 4.8–5.6× band. At the other extreme, sum ties the CPU at 1.03×: AVX2 auto-vectorization on the EPYC reaches full memory bandwidth at the 1Melement size, erasing the GPU’s edge for this bandwidth-bound reduction. The behaviour at small sizes is governed by a fixed GPU launch floor of roughly 35–36 µs. Below input sizes of a few thousand elements this overhead dominates and the CPU wins; the crossover point sits between 1K and 64K elements, above which the GPU wins cleanly. This is not a defect but a design boundary: it tells an analysis framework where it is worth dispatching to the accelerator and where it is not. It is also why the argsort kernel is a striking case—the GPU beats all 128 CPU cores even on relatively small lists, and even with the PCIe copy included, because the sorting work per element is high enough to amortize both the launch floor and the transfer.
reduce_sum_i64 reduce_prod_i64 compact_offsets reduce_max_i64 reduce_sum_bool reduce_min_i64 reduce_argmax reduce_argmin countnonzero reduce_count missing_repeat reduce_prod_bool
0.08 0.08 0.20 0.21 0.22 0.22 0.66 0.66 0.95 1.00 1.00 1.08
0.0 0.2 0.4 0.6 0.8 1.0 1.2 Runtime ratio Rust / C++ (geomean; < 1 = Rust faster) Figure 2. Rust-versus-C++ CPU kernel runtime ratio (geometric mean over three input sizes; lower is better, dashed line marks parity). Green bars are Rust wins, grey is parity, red is a regression. Overall geometric-mean ratio 0.37× across twelve kernels.
8.2 Rust versus C++ CPU kernels
The foundry model requires that the Rust kernels be at least as good as the C++ kernels they are meant to replace, on the CPU baseline, before any GPU discussion is meaningful. Figure 2 reports the per-kernel runtime ratio of the Rust implementation to the incumbent C++ implementation (Criterion mean over 100 samples; values below 1.0 mean Rust is faster). The Rust kernels are dramatically faster on the reductions that vectorize well (reduce_sum_i64 at 82.9 µs versus 1.64 ms, and reduce_prod_i64, both near 0.08×—an order of magnitude), comfortably faster on offset compaction and min/max (∼0.20×), and at parity on the remaining kernels, with a single kernel (prod_bool) marginally slower at the smallest size but recovering to parity at scale. The overall geometric-mean runtime ratio is 0.37× across twelve kernels: the Rust core is, on aggregate, comfortably faster than the C++ baseline while adding compile-time memory-safety guarantees. 8.3 Profiling summary
Taken together, the measurements draw a consistent picture: a hard GPU launch floor near 35 µs sets the small-size crossover; write-dominated kernels (count at 12.5×, countnonzero at 9.5×) extract the largest speedups because they are limited only by memory throughput; compare-reduce kernels (argmax/argmin/max/min/prod) occupy a 4.8–5.6× middle tier; and bandwidth-bound reductions such as sum reach only parity once AVX2 saturates CPU bandwidth. The lesson for a portable engine is that the kernel’s memory-access character, not the vendor label, predicts where acceleration pays off—once the HIP-specific patterns of Sect. 6 have removed the artificial penalty of a naïve port.
9 Roadmap and future work Four directions structure the work ahead. Kernel completeness: extend beyond reductions so that carry, flatten, mask and sort join reduce across every numeric type on CPU, Metal and HIP. Adaptive performance: replace hand-tuned launch constants with profile-guided blockand tile-selection chosen at first run per kernel and per GPU architecture and then cached, with HIP auto-tuning validated on MI300A. Unified GPU backend: collapse the current
per-target kernel code behind a single GpuBackend trait spanning Metal, HIP and CUDA, so there is one dispatch path and zero per-target boilerplate in kernel code—the macro dispatch of Sect. 4 is the first step toward this. Ecosystem integration: provide native columnar interoperability with RDataFrame and coffea so that rawkward arrays are usable in HEP pipelines without copy or conversion, and validate the whole stack on MI300A-class clusters. Looking further ahead, the CDNA4 MI350 generation and its 288 GB of HBM3E memory will relax the memory-capacity constraints that currently shape kernel tiling.
10 Conclusion HIP kernels for irregular data are not trivial: a direct port of an optimized CUDA kernel can run 5–10× slower, and the reason is architectural rather than a matter of toolchain maturity. But the gap is closable. With a small, reusable set of patterns—loop flattening, 128-bit vectorized loads, splitting kernels rather than fusing them, and a Rust macro that dispatches each call to the right vendor backend—the HIP kernels match their CUDA counterparts and unlock real performance portability. On a two-socket MI210 node the resulting backend delivers up to 12.5× over 128 CPU cores where the memory pattern favours the GPU, and reaches parity precisely where the CPU already saturates its bandwidth, while the underlying Rust kernels beat the incumbent C++ kernels by a geometric-mean factor of 0.37× on the CPU baseline. The same Rust API drives both vendors, with the hardware-specific choices hidden behind the dispatch layer. For the HL-LHC era, in which AMD accelerators are no longer the exception, that combination—one API, parity performance, memory safety by construction—is a practical route toward a truly hardware-agnostic HEP analysis stack.
Acknowledgements This work was supported by the U.S. National Science Foundation (NSF) under Cooperative Agreements OAC-1450377, OAC-1836650, OAC-2103945, PHY-2121686 and PHY2323298, and carried out within the Institute for Research and Innovation in Software for High Energy Physics (IRIS-HEP). The authors thank the IRIS-HEP and Awkward Array communities, and acknowledge the Princeton Research Computing della-milan resources used for the benchmarks.
References [1] J. Letts, WLCG Technical Evolution: Preparing for HL-LHC, these proceedings [2] J. Rembser, Particle physics data analysis on the GPU, these proceedings [3] K.A. Mohrman, GPU acceleration of end-user analyses at the LHC, these proceedings [4] J. Pivarski et al., Awkward Array, Zenodo (2018), https://github.com/scikit-hep/ awkward [5] M. Naumchyk et al., CUDA Acceleration of Awkward Array Using Python CCCL, these proceedings [6] M. Naumchyk et al., Coffea schema modifications for GPU backends, these proceedings [7] I. Osborne, Bridging the Vendor Gap: benchmark data and scripts, GitHub (2026), https://github.com/ianna/CHEP2026-Bridging-the-Vendor-Gap [8] Advanced Micro Devices, ROCm/HIP Programming Guide (2025), https://rocm.docs. amd.com [9] Advanced Micro Devices, AMD CDNA2 Architecture White Paper (2022) [10] B. Heisler et al., Criterion.rs: Statistics-driven benchmarking for Rust (2024), https: //github.com/bheisler/criterion.rs