Conceptio › Archive › arXiv CS
arXiv CSopen access

Towards Standardized Evaluation of GPU Memory Safety with GMSBench

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

1

Towards Standardized Evaluation of GPU Memory Safety with GMSBench

arXiv:2609.08871v1 [cs.CR] 8 Sep 2026

Saurabh Singh∗ , Jaewon Lee† , Seonjin Na‡ , and Hyesoon Kim∗ ∗ Georgia Institute of Technology, † Microsoft, ‡ NVIDIA [email protected], [email protected], [email protected], [email protected]

Abstract—As GPUs become increasingly integral to highperformance computing and machine learning, ensuring memory safety in GPU programs has become crucial for reliable and secure execution. However, evaluating GPU memory safety techniques remains challenging due to the lack of comprehensive and standardized benchmarks. In this paper, we present GMSBench, a GPU memory safety benchmark designed to evaluate a broad range of memory safety violations across different GPU memory spaces and execution scenarios. GMSBench comprises 149 selfcontained CUDA tests spanning spatial, temporal, and concurrency errors. The suite provides a standardized foundation for the evaluation and comparative analysis of GPU memory safety mechanisms and helps expose gaps in their detection coverage. We demonstrate the utility of GMSBench by evaluating Compute Sanitizer, a widely used GPU memory error detection tool across multiple GPU architectures.

I. I NTRODUCTION Graphics Processing Units (GPUs) have become a cornerstone of modern computing, driving advancements in fields such as scientific computing, machine learning, and graphics rendering. The massively parallel architectures make them highly effective at handling intensive parallel workloads, but also introduce unique challenges for ensuring memory safety. In GPU applications, memory safety violations such as buffer overflows, out-of-bounds accesses, use-after-free, or race conditions can lead to erroneous results, crashes, or even security vulnerabilities. Recent studies have shown that GPU memory safety errors can be exploited to develop attacks that degrade application performance [1], hijack control flow [2], and even cross the CPU-GPU isolation layer and corrupt CPU memory [3]. Despite growing interest in GPU memory safety, evaluation of memory error detection and protection mechanisms remains challenging. GPU programs operate across a heterogeneous memory hierarchy that includes global, shared, local, and constant memory, each with distinct access, visibility, and sharing semantics. Existing works often rely on tool-specific collections of test cases, with substantial variation in the types of errors, memory spaces, and access patterns they cover. As a result, reported detection coverage is difficult to compare across techniques, and important classes of violations may remain untested. This lack of a common and systematic benchmark makes it harder to assess the strengths and limitations of existing and future GPU memory safety mechanisms. Prior GPU memory safety works have evaluated their techniques using custom collections of memory error tests. For

example, cuCatch [4] evaluates its approach using a closedsource benchmark suite, making it difficult to inspect the covered scenarios or reproduce the evaluation. In contrast, CuSafe [5] provides an open-source set of tests, but the suite is relatively small and covers only a limited subset of GPU memory safety errors. More broadly, existing benchmarks differ substantially in availability, organization, documentation, and coverage, making it difficult to determine which vulnerability classes, memory spaces, and access patterns are exercised and to directly compare results across techniques. To address this gap, we present GMSBench, an opensource benchmark suite designed to provide a common and systematic foundation for evaluating hardware- and softwarebased GPU memory safety mechanisms. GMSBench contains carefully designed tests spanning across a broad range of vulnerability classes, memory spaces, allocation mechanisms, and access behaviors, enabling comprehensive and reproducible evaluations. This broad coverage allows GMSBench to expose weaknesses and blind spots in existing memory safety mechanisms while providing a consistent basis for comparison across different approaches. Overall, our contributions are as follows: 1) We develop a systematic classification of GPU memory safety violations across spatial, temporal, and concurrency-related errors, and identify key dimensions affecting their manifestation and detection. 2) We present GMSBench, an open-source1 benchmark suite that builds on this classification to evaluate GPU memory safety mechanisms. 3) We demonstrate the utility of GMSBench by evaluating a widely used GPU memory error detection tool across multiple GPU architectures and generations, and characterize its detection coverage across the benchmark. II. BACKGROUND AND M OTIVATION A. GPU Memory Safety GPUs expose a heterogeneous memory hierarchy with distinct address spaces, allocation mechanisms, lifetimes, and visibility rules. CUDA programs may access GPU global memory allocated by the host or device, per-thread local memory, per-block shared memory, constant memory, and host-pinned memory. These memories differ in how objects are allocated, addressed, and shared among GPU threads. 1 Available at: https://github.com/gthparch/GMSBench

2

As a result, memory-safety violations on GPUs can manifest differently from those in conventional CPU programs. For example, an out-of-bounds access may cross directly from one valid global-memory allocation into another without accessing unmapped memory in between. Similarly, sharedmemory violations may corrupt neighboring regions within a thread block, while local-memory violations depend on per-thread stack organization and object lifetime. Pinnedmemory violations can directly affect data resident in host RAM, extending the impact of GPU memory errors beyond device memory. GPU programs also introduce concurrencyrelated errors whose manifestation depends on the relationship between threads, warps, and thread blocks. These architectural characteristics make GPU memory safety a multidimensional problem involving not only the type of violation, but also the affected memory space, allocation mechanism, object layout, and execution scope. B. Challenges in Evaluating GPU Memory Safety A variety of hardware and software mechanisms have been proposed to detect or prevent GPU memory-safety violations [4]–[8]. However, evaluating these mechanisms remains difficult because existing studies rely on a small set of tests often tailored for a specific technique. Although such tests may demonstrate detection of individual bugs, they provide limited insight into overall coverage of the mechanism. Detection behavior can vary across violation types, memory spaces, allocation mechanisms, object layouts, and execution scopes. Existing mechanisms use a variety of detection strategies, such as allocation-bound tracking [4], [5], padding/redzones [6], and/or runtime metadata, and therefore expose different blind spots. For example, two global memory overflows may be detected differently based on whether the victim object is adjacent to the overflowing object or separated by an unallocated region or a live object. Likewise, mechanisms that detect global-memory violations may not provide equivalent coverage for local or shared memory, and race-detection capabilities may depend on whether racing threads reside within the same warp, block, or different blocks. A useful benchmark must therefore expose these dimensions explicitly while isolating individual violations so that each result can be attributed to a specific detection capability. The test cases should also remain independent of detection mechanisms, allowing the same violations to be evaluated consistently across tools. Another challenge is establishing a reliable ground truth. A crash or non-zero return code does not necessarily indicate that the intended violation was detected, since the failure may result from a runtime error or a secondary effect of memory corruption. Similarly, an invalid access may silently corrupt another valid object without causing a crash. In some cases, a violation may be semantically present in the program but may not manifest as an observable error during native GPU execution. Evaluation must therefore distinguish between whether the violation actually manifested and whether the mechanism detected it. This separation is essential for accurately characterizing detection coverage and comparing different GPU memory-safety mechanisms. These challenges motivate a benchmark that systematically captures GPU-specific

memory-safety behaviors while independently establishing the manifestation of each violation whenever possible. III. GPU M EMORY S AFETY B ENCHMARK Building on the challenges identified in Section II-B, we present GMSBench, a GPU memory safety benchmark suite inspired by prior efforts [4], [5]. GMSBench comprises 149 self-contained CUDA tests covering spatial errors, temporal errors, and data races across various GPU memory spaces and object relationships, as summarized in Figure 1. A. Design Principles GMSBench is designed around four key principles: 1) Systematic Coverage: The benchmark systematically covers the dimensions that influence the manifestation and detection of GPU memory-safety violations rather than exhaustively enumerating all possible combinations. GMSBench includes variants only when they represent a semantic distinction or can expose different detector behavior in current and future mechanisms. 2) Violation Isolation: Each test is self-contained and isolates a single memory safety violation, allowing the results to be attributed to a specific detection capability. 3) Independent Ground Truth: Whenever possible, GMSBench establishes ground truth through baseline execution that runs each test natively on the GPU without any protection mechanism. The baseline run checks for an observable effect such as sentinel corruption, disclosure of a planted value, or an incorrect result to determine whether the violation actually manifested. Tests that do not show reliable native manifestations are marked as detection-only. 4) Mechanism-Agnostic: Tests remain independent of any particular detection mechanism. The shared GMSBench harness provides common allocation, layout, and ground-truth helpers, allowing the same test logic to be evaluated across different mechanisms. B. Classification of GPU Memory Safety Violations 1) Taxonomy: GMSBench organizes tests along four dimensions as shown in Figure 1. 1 Error class identifies whether the violation is spatial, temporal, or a data race. 2 Memory window identifies the memory domain(s) involved, such as global, local, shared, or constant memory, including cross-space accesses. 3 Object context captures how participating objects are allocated and how they relate to one another. This includes allocation mechanisms in global memory, static or dynamic placement in shared memory, stackframe placement in local memory, and intra-object placement. 4 Violation type identifies the specific memory-safety error being exercised, such as linear buffer over-reads/writes, targeted out-of-bounds read/write, use-after-free (UAF), invalid free (IF), double free (DF), uninitialized read (UR), use-afterscope (UAS), use-after-return (UAR), and lost update (LU). Additional class-specific variants, such as victim adjacency, access direction, or reuse timing, are included where they represent meaningful differences in behavior or detection.

3

2) Spatial Memory Safety: Spatial errors occur when a GPU thread accesses memory outside the intended bounds of an object. GMSBench covers spatial violations across global, local, shared, and constant memory. It also includes a crossspace category where a pointer derived from one memory space is used to access an object in another. It distinguishes linear buffer over-reads/writes from targeted out-of-bounds reads/writes and includes additional object-context variants where they may affect detector behavior. Global memory tests cover device-memory allocations created using host-side cudaMalloc(), device-side in-kernel malloc(), and pinned host-memory allocations created using cudaHostAlloc() that are accessible to GPU kernels.2 Global memory tests vary the relationship between the source and victim objects, including adjacent objects and objects separated by a freed gap or another live object. Local-memory tests distinguish accesses within a single stack frame from accesses that cross function-frame boundaries, and cover both statically and dynamically allocated stack objects. Shared memory tests cover interactions between static and dynamically allocated shared memory objects within a thread block. The adjacent and non-adjacent object relationships are included only in the static-to-static case where a live intervening object can be constructed reliably. Constant memory tests include linear and targeted out-of-bounds read variants, since constant memory cannot be written from device code. GMSBench also includes intra-object violations in global, local, shared, and constant memory, where a linear or targeted access crosses field boundaries within a single object. These tests probe whether detectors can distinguish sub-

object boundaries instead of relying only on allocation-level bounds. Cross-space tests use pointer arithmetic to redirect accesses from one GPU memory space to another.3 These tests probe whether the detector distinguishes and tracks pointer provenance across GPU memory spaces rather than validating only the final address. Finally, since local and shared memory occupy bounded windows, GMSBench includes a beyondwindow category in which a linear access reaches beyond the entire memory window. 3) Temporal Memory Safety: Temporal errors involve accesses or deallocations whose validity depends on an object’s lifetime or initialization state. GMSBench covers use-after-free (UAF), invalid free (IF), double free (DF), and uninitialized read (UR) errors in global memory, as well as use-after-return (UAR) and use-after-scope (UAS) errors in local memory. UAF, IF, and DF are not included for shared memory since shared memory is scoped to a thread block and has no explicit object deallocation during kernel execution. Furthermore, a reliable ground truth could not be established for sharedmemory uninitialized reads. For global memory, GMSBench exercises temporal violations across device-memory allocations created using hostside cudaMalloc(), device-side in-kernel malloc() allocations, managed-memory allocations, and pinned hostmemory allocations. UAF tests distinguish between immediate accesses, where a dangling pointer is used to access a freed region after deallocation, and delayed accesses, where the same address has been reallocated to a different object before the dangling access. This distinction probes whether a detector recognizes only memory that is currently invalid or also tracks allocation identity after the same address is reused. Aliased UAF tests free an object through one pointer and later access it through a copied alias, probing whether detectors track object lifetime across pointer aliases. Managed memory contributes only a delayed UAF variant because a freed managed region is unmapped on deallocation, causing a dangling read to fault rather than disclose the planted value, unlike cudaMalloc() memory, which remains readable after deallocation. IF tests exercise deallocation using pointers that do not point to the start of a valid allocation or that originate from mismatched allocators. DF tests probe deallocation of the same object more than once. UR tests probe reads that occur before the object’s contents have been initialized. Local-memory tests also cover two stack lifetime errors: use-after-scope (UAS) and use-after-return (UAR). UAS tests access local objects after their lexical scope ends, while UAR tests access local objects after the defining function returns. These tests also include immediate and delayed variants depending on whether the expired stack object is accessed before or after subsequent reuse. These cases test whether detectors track GPU localmemory lifetimes across scopes and function calls. 4) Data Races: Data races occur when multiple GPU threads access the same memory location without synchronization, with at least one write. In GMSBench, data-race tests use unsynchronized non-atomic read-modify-write operations

2 cudaMallocManaged() contributes no spatial tests because its tested behavior was equivalent to cudaMalloc().

3 Linear overflows are excluded because they reach the source memory window boundary before reaching another memory window

Spatial Global

Constant *

Local

Host allocated

Device allocated

Host pinned

Intra-obj

Intraframe

Interframe

St

St

Dy

Beyond

Intra- Interobj obj

Shared St to St

St to Dy

G to L

G to S

Dy

Dy to St

Dy to Dy

L to G

L to S

Intra-obj

Beyond

Intra-obj

S to G

S to L

Temporal Global

Race Local

Host allocated

Device allocated

Host pinned

Managed

Cross-space

Global

In warp

In block

Shared

In warp

In block

Cross block

1

Error class

Linear over-read/write

Uninitialized read

2

Memory window

Targeted OOB R/W

Use-after-scope

3

Object context

Use-after-free

Use-after-return

43

Violation type

Invalid free

Lost update

Double free * Read only accesses. G: Global, L: Local, S: Shared, St: Static, Dy: Dynamic Beyond: access leaves the local/shared memory window.

Fig. 1. Classification of GPU Memory-Safety Violations in GMSBench

4

on a common counter, causing lost updates that provide a deterministic observable effect. The benchmark varies the relationship between racing threads; shared-memory races cover intra-warp and intra-block cases, while global-memory races additionally include cross-block cases. These tests evaluate whether detectors can identify data races across different GPU thread relationships and memory scopes. C. Test Design Every test is a self-contained CUDA file defining exactly one gms_test() function. A common harness provides the program entry point, allocation and layout helpers, allowing individual tests to focus on the violation being exercised. As shown in Listing 1, a typical test follows three steps: (1) construct the memory layout required by the test, (2) trigger exactly one memory safety violation, and (3) expose an observable effect that can be checked during baseline execution.

GMSBench uses a config-driven runner to execute the tests across different targets. Each target specifies how a test is compiled and executed, along with diagnostic patterns used to identify tool-reported detections. This separates benchmark ground truth from mechanism-specific detection and allows the same tests to be evaluated consistently across targets. If a required layout or allocation condition cannot be established, the runner reports SETUP_FAIL rather than treating it as a detector miss. IV. E VALUATION

We evaluate GMSBench by measuring (1) violation manifestation under native CUDA execution and (2) detection coverage using NVIDIA Compute Sanitizer (CompSan) [6], with its memcheck, initcheck, and racecheck tools. We also include a variant of memcheck with API errors enabled and call it apicheck.4 Each tool runs on the applicable set of GMSBench tests. To assess the robustness of the benchmark, 1 #include "gms.h" we evaluate the suite on three NVIDIA GPU platforms: 2 3 __global__ void test_kern(uint8_t *buf, uint8_t *sentinel) { a workstation-grade RTX A5000 (Ampere, SM_86) and a 4 uintptr_t dist = (uintptr_t)sentinel - (uintptr_t)buf; datacenter A100 (Ampere, SM_80), both with x86 hosts, and 5 for (size_t i = 0; i < dist + 4; i++) 6 buf[i] = 0xef; // overflow a datacenter GH200 Grace Hopper system with an AArch64 7 } Grace CPU and Hopper GPU (SM_90). 8 9

10 11 12 13 14 15 16 17 18 19 20 21 22 23

void gms_test() { uint8_t h_expected[16], h_got[16]; uint8_t *d_buf, *d_sentinel; init_sentinel(h_expected, 16); // Allocate adjacent attacker and victim buffers gms_cuda_malloc_adjacent(&d_buf, &d_sentinel, 32, 16); cudaMemcpy(d_sentinel, h_expected, 16, cudaMemcpyHostToDevice); // Trigger violation test_kern<<<1, 1>>>(d_buf, d_sentinel); // Detect sentinel corruption cudaMemcpy(h_got, d_sentinel, 16, cudaMemcpyDeviceToHost); CUDA_CHECK_LAST_ERROR(); gms_check_sentinel(h_got, h_expected, 16); gms_print_test_result(); }

Listing 1. Adjacent-buffer overflow test in GMSBench.

The layout helpers are used to construct controlled memory layouts rather than assuming a particular allocator behavior. For example, gms_cuda_malloc_adjacent repeatedly allocates buffers and measures their addresses until the required attacker-victim placement is achieved. Similar helpers are used to create non-adjacent layouts, pinned-memory layouts, and reclaimed allocations for temporal tests. We establish ground truth by running each test natively on the GPU without any protection mechanism. For write violations, GMSBench checks whether a planted sentinel value was corrupted; for OOB reads, it checks whether a planted value was disclosed; and for data races, it checks whether execution produced a deterministic incorrect result. For invalid-free and double-free tests with a reliable runtime signal, GMSBench uses the CUDA runtime’s rejection of the illegal deallocation as the observable effect. For uninitialized reads, GMSBench checks whether a known stale value is observed when reliable native manifestation is possible. For violations that are semantically present but do not produce a reliable native manifestation, GMSBench marks the test as detection-only and evaluates the mechanism using its own diagnostic output.

TABLE I GMSB ENCH R ESULTS ACROSS GPU P LATFORMS Test Classification

No. tests

host-allocated device-allocated host-pinned intra-object intra-frame inter-frame local intra-object beyond static2static static2dynamic dynamic2static spatial shared dynamic2dynamic intra-object beyond inter-object constant intra-object global2local global2shared local2global cross shared2global local2shared shared2local host-allocated device-allocated global temporal managed host-pinned local in warp global in blk race cross block in warp shared in blk

10 10 10 4 16 8 4 2 8 4 4 4 4 2 4 2 2 2 2 2 2 2 13 6 2 7 8 1 1 1 1 1 149

global

Totals

RTX A5000 / A100 / GH200 native memcheck apicheck initcheck racecheck 10 2 10 6 10 2 4 0 16 0 8 0 4 0 ∗ 0 2 8 0 4 0 4 0 4 0 4 0 ∗ 0 2 4 0 2 0 2 0 2 0 2 0 2 0 2 0 2 0 ∗ 12 1 5 (5) 3 (3) ∗ 4 2 1 (2) 1 (1) 2 0 7 1 4 (4) ∗ 0 0 1 0 1 0 1 0 1 1 1 1 134/149 18/144 10/11 4/4 2/5 (90%) (12.5%) (91%) (100%) (40%)

Native: number of tests whose intended violation manifests during native CUDA execution. ∗ One or more tests do not manifest during native execution and are included for detection only x(y) means x tests were detected among y applicable tests

Cross-platform invariance: We ran the benchmarks compiled with the -arch=native flag so that device code is generated for the GPU it runs on. As shown in Table I, the native manifestation and tool verdicts are identical across all GPU architectures we tested. The invariance holds even for the local-memory tests, where register allocation and stack layout 4 memcheck with --report-api-errors=explicit

5

genuinely differ across the three targets. This is because the suite assigns attacker and victim roles by measured address at runtime rather than by declaration order. The invariance also holds for data races despite their schedule-dependent execution. Compute Sanitizer Coverage: In our experiments, CompSan variants behaved consistently across all platforms. Memcheck, CompSan’s memory-error detection tool, detected only 18/144 tests across the spatial and temporal error classes. For global memory, memcheck detects linear overflows that cross a freed gap but misses overflows into an adjacent buffer or a nonadjacent buffer separated by a live intervening object. Since memcheck uses a tripwire-based mechanism [4], it misses linear overflows which land directly in another live allocation without crossing an invalid region. It misses targeted OOB accesses that jump over the tripwire and land in another valid memory region. It also misses all intra-object and cross-space violations, indicating that the tool does not track sub-object boundaries or pointer provenance. An important exception is device-heap memory allocated using in-kernel malloc(), where memcheck detects all linear overflow variants, including accesses into live neighboring objects. For temporal errors, memcheck detects immediate UAF accesses but misses delayed and aliased UAFs. This indicates that memcheck does not track allocation identity after an address is reused by another object. Other CompSan tools like apicheck, racecheck, initcheck provide stronger coverage within their intended scopes. apicheck detects 10/11 invalid- and double-free tests by reporting CUDA API errors, but misses the device-heap double free performed through in-kernel free(). initcheck detects all 4/4 uninitialized-read tests. racecheck detects 2/5 race tests, correctly identifying both shared-memory races while missing all three global-memory races. V. R ELATED W ORK Prior work has developed synthetic test suites to evaluate GPU memory-safety tools. cuCatch [4] used a set of 56 CUDA tests covering spatial and temporal violations across global, local, and shared memory but is not publicly available and uses a relatively coarse classification. It also does not cover dimensions such as cross-memory-space violations and data races. CuSafe [5] provides an open-source benchmark containing 33 tests covering spatial and temporal errors, including linear and non-linear overflows and local memory violations, but remains limited in size and taxonomy. Neither suite systematically covers constant memory, host-pinned memory, cross-window provenance violations, uninitialized reads, aliased UAF, or data races under a single taxonomy. In contrast, GMSBench provides a larger and more fine-grained benchmark. Table II summarizes the differences in benchmark coverage based on the taxonomy defined in Section III-B1. Beyond benchmarks, several works have proposed GPU memory-safety tools for CUDA and OpenCL with different detection models. These use a range of detection mechanisms, including redzones/canaries, object-bounds tracking, pointer metadata, compiler instrumentation, runtime checking, and hardware-software co-design [4], [5], [7]–[10]. These

approaches provide different coverage across memory spaces, allocation mechanisms, and error classes, motivating a common benchmark for systematic comparison. TABLE II C OMPARISON OF E XISTING GPU M EMORY-S AFETY B ENCHMARKS Name CuSafe [5] cuCatch [4] GMSBench (Ours)

No. Tests 33 56 149

Opena ✓ ✗ ✓

Error Class

Memory Window

Object Context Violation Type

Circle fill denotes coverage of the corresponding GMSBench taxonomy a Benchmark source-code availability

VI. CONCLUSION In this paper, we presented a systematic taxonomy for GPU memory safety and introduced GMSBench, an opensource benchmark suite for systematic evaluation of GPU memory safety mechanisms. GMSBench covers diverse GPU memory spaces, allocation mechanisms, access patterns, and execution scopes, providing a common foundation for reproducible evaluation and comparison. Our evaluation across three GPU platforms shows consistent benchmark behavior and reveals several systematic gaps in Compute Sanitizer’s detection coverage. R EFERENCES [1] S.-O. Park, O. Kwon, Y. Kim, S. K. Cha, and H. Yoon, “Mind control attack: Undermining deep learning with gpu memory exploitation,” Comput. Secur., vol. 102, no. C, Mar. 2021. [Online]. Available: https://doi.org/10.1016/j.cose.2020.102115 [2] Y. Guo, Z. Zhang, and J. Yang, “GPU memory exploitation for fun and profit,” in 33rd USENIX Security Symposium (USENIX Security 24). Philadelphia, PA: USENIX Association, Aug. 2024, pp. 4033–4050. [Online]. Available: https://www.usenix.org/conference/ usenixsecurity24/presentation/guo-yanan [3] S. Roh, W. Choi, J. Chung, Y. Lee, S. Song, and B. Lee, “Ghost in the shell: A gpu-to-host memory attack and its mitigation,” in 2026 IEEE Symposium on Security and Privacy (SP), 2026, pp. 456–471. [4] M. Tarek Ibn Ziad, S. Damani, A. Jaleel, S. W. Keckler, and M. Stephenson, “cucatch: A debugging tool for efficiently catching memory safety violations in cuda applications,” Proc. ACM Program. Lang., vol. 7, no. PLDI, Jun. 2023. [Online]. Available: https://doi.org/10.1145/3591225 [5] H. Lu, F. Zhang, Z. Zhang, S. Wang, and Y. Guo, “CuSafe: Capturing memory corruption on NVIDIA GPUs,” in 35th USENIX Security Symposium (USENIX Security 26). Baltimore, MD: USENIX Association, Aug. 2026. [Online]. Available: https://www.usenix.org/ conference/usenixsecurity26/presentation/lu-hongyi [6] “Nvidia compute sanitizer,” 2026. [Online]. Available: https://developer. nvidia.com/compute-sanitizer [7] J. Lee, Y. Kim, J. Cao, E. Kim, J. Lee, and H. Kim, “Securing gpu via region-based bounds checking,” in Proceedings of the 49th Annual International Symposium on Computer Architecture, ser. ISCA ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 27–41. [Online]. Available: https://doi.org/10.1145/3470496.3527420 [8] J. Lee, E. Chung, S. Singh, S. Na, Y. Kim, J. Lee, and H. Kim, “Letme-in: (still) employing in-pointer bounds metadata for fine-grained gpu memory safety,” in 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2025, pp. 1648–1661. [9] M. Tarek Ibn Ziad, S. Damani, M. Stephenson, S. W. Keckler, and A. Jaleel, “Gpuarmor: A hardware-software co-design for efficient and scalable memory safety on gpus,” ACM Trans. Archit. Code Optim., vol. 23, no. 3, Jul. 2026. [Online]. Available: https://doi.org/10.1145/3815783 [10] B. Di, J. Sun, D. Li, H. Chen, and Z. Quan, “Gmod: a dynamic gpu memory overflow detector,” in Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques, ser. PACT ’18. New York, NY, USA: Association for Computing Machinery, 2018. [Online]. Available: https://doi.org/10.1145/3243176. 3243194

Record · ID 667888 · SHA-256 87bdc4f4a2f641af
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.