Xiaoze Fan
Jianhao Wang
Weihao Cui
[email protected] Shanghai Jiao Tong University
[email protected] Shanghai Jiao Tong University
[email protected] Shanghai Jiao Tong University
Han Zhao
Zhuobin Huang
Yangjie Zhou
[email protected] Shanghai Jiao Tong University
[email protected] National University of Singapore
[email protected] National University of Singapore
Yuxian Qiu
Shixuan Sun
Bingsheng He
[email protected] NVIDIA
[email protected] Shanghai Jiao Tong University
[email protected] National University of Singapore
Quan Chen
Minyi Guo
[email protected] Shanghai Jiao Tong University
[email protected] Shanghai Jiao Tong University Floorswept SMs
Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning. The former creates chipspecific compute topologies, while the latter causes nonuniform memory access. These asymmetries are substantial. Topology-oblivious compute unit allocation can lead to up to 1.33× performance variation, while remote accesses increase HBM latency by up to 67% and nearly double L2 latency. However, these asymmetries are hidden behind the GPU’s logical resource abstractions and can vary across chips. We develop lightweight characterization methods to uncover per-chip compute topology and memory affinity. We then use the discovered information to make existing fine-grained scheduling asymmetry-aware, considering not only how many resources are allocated but also which physical resources are assigned. Across full-GPU kernel execution, intraapplication multiplexing, and inter-application co-location, asymmetry-aware scheduling improves mainstream kernels by up to 1.22×, multiplexed LLM inference by up to 14.3%, and avoids up to 1.33× performance variation.
1
TPC TPC TPC TPC TPC TPC TPC TPC TPC SM SM
SM SM
SM SM
GPC GPC GPC GPC
SM SM
SM SM
SM SM
SM SM
SM SM
Xbar LTC fabric
L2 Cache
Xbar
Xbar
GPC GPC GPC GPC
GPC GPC GPC GPC
Warp Scheduler Register File
FP32 FP64
GPC GPC GPC GPC
Xbar
L2 Cache
SM SM
INT32
HBM
Abstract
HBM
arXiv:2609.24270v1 [cs.AR] 21 Sep 2026
Dissecting How Die Scaling Breaks GPU Fine-grained Scheduling
Tensor Core
Tensor Memory Accelerator L1 Cache / Shared memory
Figure 1. Architecture of an NVIDIA H200 GPU. Streaming Multiprocessors (SMs) are grouped into Texture Processing Clusters (TPCs), which in turn form Graphics Processing Clusters (GPCs). Each SM has a private L1 cache, while all SMs share a partitioned L2 cache backed by HBM. Existing fine-grained scheduling works [1–4, 6, 8–10] generally operate by controlling logical resource quantities. For example, they control how many compute units a task uses and how much memory it accesses. These works omit the physical locations of compute units and memory. It implicitly assumes architectural symmetry, that allocations with the same number of compute units and the same memory capacity offer equivalent performance. Die scaling introduces asymmetries, which increases the number of transistors per die by enlarging die area [11] and shrinking process nodes (7 nm→3 nm) [12]. Figure 1 shows a modern NVIDIA GPU architecture and how die scaling affects it. First, floorsweeping disables defective Streaming Multiprocessors (SMs, NVIDIA GPUs’ basic compute units) to improve yield, creating chip-specific compute topologies. Second, larger L2 caches are partitioned for high bandwidth and low latency. These partitions and their associated HBM form a non-uniform memory access (NUMA) topology with different local and remote access costs. These asymmetries have substantial performance consequences. Floorsweeping creates compute asymmetry. Two
Introduction
Modern GPUs integrate many parallel resources. Efficiently utilizing these resources requires fine-grained scheduling at multiple levels. Kernel libraries schedule work across compute units to approach hardware performance limits [1, 2]. Within an application, recent LLM serving systems spatially multiplex prefill and decode phases to improve goodput [3, 4]. Across applications, cloud systems co-locate kernels from different tenants on the same GPU and control each tenant’s compute and memory allocation [5–10]. 1
Xiaoze Fan et al.
Table 1. GPU die specifications across generations. H100 ships in SXM and PCIe variants; this table lists the PCIe SKU. H100 PCIe and H200 share the GH100 die but differ in enabled resources. B200 fuses two GB100 chiplets; values marked ×2 denote per-chiplet quantities, while unmarked values are package totals.
GPCs can differ by up to 10 SMs on either NVIDIA H200 or B200. Remote HBM accesses incur approximately 34% higher latency than local accesses on H200 (∼490→∼655 cycles) and 67% on B200 (∼552→∼920 cycles). Remote L2 latency is about 51% higher than local latency on H200 (∼309→∼466 cycles) and nearly twice as high on B200 (∼364→∼725 cycles). As we show later in §7, ignoring these differences can cause up to 1.33× performance variation. This paper addresses two challenges raised by such hidden physical asymmetry. C-1: How can software efficiently discover hidden and chip-specific compute and memory asymmetries without architectural documentation? Vendor exposes logical rather than physical SM topology, floorsweeping differs across chips of the same SKU1 , NUMA mapping is encoded in undocumented physical-address hashing, and different architectures use different mappings. Thus, asymmetry often requires per-chip calibration rather than a static lookup table shared across chips of the same SKU. C-2: How should existing fine-grained scheduling exploit the discovered physical asymmetry? Existing scheduling works determine only resource quantity. Asymmetry awareness adds a complementary dimension by considering physical resource identity, topology, and compute-memory affinity. The goal is not to replace fine-grained scheduling, but to make it more precise by considering not only how many resources are allocated, but which physical resources are allocated. To address the first challenge, we develop lightweight characterization methods for compute and memory asymmetry. For compute asymmetry, we uncover the complete per-chip SM-to-GPC assignment using two probing techniques. We observe two groups of logical SM IDs. Normal SMs follow an architecture-specific predefined SM-to-GPC mapping, whereas the remaining SMs are assigned to physical GPCs depending on the chip-specific floorsweeping outcome. For convenience, we refer to the latter as random SMs. “random” does not mean runtime-random scheduling, but that their physical GPC locations may vary across chips. For memory asymmetry, we develop a latency hierarchy discovery method that uncovers the memory partition hash and HBM interleaving granularity on the measured GPUs. With the uncovered hash, we characterize local/remote latency gaps, remote bandwidth bottlenecks, and effective L2 capacity loss from cross-partition cache replication. To address the second challenge, we build asymmetryaware prototypes for three representative fine-grained scheduling scenarios. For full-GPU kernels, we develop two methods that steers memory accesses to NUMA-local partitions. Integrating these methods with mainstream GPU kernels improves throughput by up to 1.22 × with minimal code changes. For intra-application multiplexing, we develop a
SKU Die
V100 A100 GV100 [15] GA100 [16]
Die (mm2 ) Full SMs SKU SMs Disabled SMs GPCs TPCs/GPC
815 84 80 4 6 7
826 128 108 20 7 8
L2 (MB) L2 partitions HBM type HBM (GB)
6 1 HBM2 32
40 2 HBM2e 80
H100 H200 GH100 [17]
B200 GB100 × 2 [18]
814 144
9
800 × 2 80 × 2 148 12 4×2 10
50 60 2 2 HBM2e HBM3e 80 141
126 2 HBM3e 90 × 2
114 30 7/8
132 12 8
kernel-transparent cross-NUMA allocation method for LLM serving systems that spatially multiplex prefill and decode phases. This improves decode throughput by up to 14.3% by restricting task-private allocations to NUMA-local partitions while sharing others across partitions. For inter-application co-location, we evaluate topology-aware SM allocation strategies and show that equal SM counts do not imply equal physical capability. Topology-oblivious allocation causes up to 1.33× throughput variation. We will open-source our code to enable reproduction of all our results. Our main contributions are as follows: • We develop lightweight methods to uncover hidden, chipspecific GPU compute and memory asymmetry, including SM-to-GPC mapping, floorsweeping-dependent topology, NUMA partition hashes, and local/remote latency and bandwidth behavior. • We show how fine-grained scheduling can incorporate physical resource identity, topology, and memory affinity rather than relying only on logical resource counts. • We validate asymmetry-aware scheduling across three levels. These results motivate making asymmetry awareness a design principle for future GPU programming.
2
Background
Workloads such as large language models [13, 14] demand increasing compute throughput and memory bandwidth, driving aggressive GPU die scaling. While Figure 1 shows the microarchitecture of a modern GPU die, Table 1 further summarizes five SKUs spanning four generations of NVIDIA data-center GPUs. From Volta to Blackwell, full-die SM counts grew from 84 to 160, L2 caches from 6 MB to 126 MB, and HBM stacks from 4 to 8. Compute. As shown in Figure 1, GPU compute is organized hierarchically (GPU - GPC - TPC - SM). As dies grow, defective SMs must be disabled to maintain manufacturing yield, a process called floorsweeping. The full GH100 die contains 8 GPCs of 9 TPCs each, totaling 8 × 9 × 2 = 144 SMs.
1 SKU (Stock Keeping Unit): a unique product identifier that encodes the GPU product name and interface, e.g., H200.
2
Floorsweeping permanently disables selected TPCs within each GPC, leaving H200 with 132 SMs across 8 GPCs. Which TPCs are disabled depends on per-die manufacturing defects. After floorsweeping, surviving TPCs are renumbered into a contiguous logical SM ID space, making the logical-tophysical mapping chip-specific and opaque to software. Memory. On the memory side, each SM has a private L1 data cache, and all SMs share a last-level L2 cache backed by off-chip HBM stacks. SMs communicate with the L2 through a crossbar (Xbar in Figure 1). Starting with Ampere, the L2 grew to 40 MB and was physically split into multiple physical partitions. On the NVIDIA GPUs studied in this paper, each GPU exposes two memory-affinity partitions. An LTC fabric connects these partitions to present a unified address space to software. Newer Blackwell GPUs, such as B200, integrate two chiplets into a single GPU. Official documentation describes one L2 partition per chiplet, with the two partitions connected by the LTC fabric. Whether the L2 within each chiplet is further partitioned remains unknown.
3
and SGDRC control resource shares and interference without incorporating chip-specific floorsweeping topology or SM-to-memory NUMA affinity into allocation decisions.
4
Motivation
NVIDIA publishes the total SM count, L2 cache size, and HBM capacity for each GPU product, but does not disclose the per-chip floorsweeping pattern or the mapping between compute units and memory partitions. Without this information, software cannot determine which physical resources a logical allocation actually receives. Understanding this relationship matters at multiple granularities. At the single-kernel level, thread blocks share SMs and memory partitions, and NUMA-unaware placement can create cross-partition bottlenecks that degrade overall kernel performance. At the single-application level, developers could overlap internal tasks on the same GPU more effectively if they knew the compute-to-memory affinity. At the multi-application level, cloud vendors need topology-aware partitioning strategies to ensure fairness across co-located applications. NVIDIA hides these variable architectural features from software to simplify the runtime and user-level programming model. For instance, since the introduction of MIG, certain SMs are silently disabled when MIG mode is active, yet no official documentation explains why. The goal of this paper is to characterize the variable, per-chip architectural asymmetries that die scaling introduces and to show how fine-grained scheduling can use that information. In the following sections, we first characterize the floorsweeping topology and NUMA affinity mapping. We then evaluate asymmetry-aware scheduling through several prototype implementations. We use H200 and B200 as our primary testbeds throughout this paper.
Related Work
Fine-grained scheduling of GPU. Existing works implement fine-grained scheduling at multiple granularities. Within a single kernel, persistent thread blocks [6, 8, 19] allow users to pin thread blocks to specific SMs. Thread Block Clusters [20], a scheduling feature introduced in Hopper, further leverages this for acceleration. For multi-task or co-located applications, several systems [8–10] schedule multiplexed kernels by controlling the number of SMs. All these systems control how many SMs are used, but none considers which SMs are assigned or their NUMA affinity. Vendor-supported scheduling. NVIDIA’s MPS (MultiProcess Service) [5] supports static SM partitioning. Green Contexts [3] provide SM-level partitions with optional GPC alignment. MIG (Multi-Instance GPU) [21] provides hardwareisolated compute and memory partitions that can align with NUMA boundaries. MPS’s MLOPart (Memory Locality Optimized Partition) [22] adds NUMA-aware partitioning for Blackwell and newer GPUs. Both MIG and MLOPart hide detailed physical compute and memory topology. Moreover, users cannot freely select specific physical SMs or choose between NUMA-local and interleaved placement for each device-memory allocation. Reverse-engineered scheduling. FGPUs [23] and SGDRC [7] reverse-engineer physical memory mappings and use coloring to reduce interference between co-located workloads. FGPUs combines page coloring with persistent blocks to reserve SMs and memory bandwidth. SGDRC dynamically allocates SMs and VRAM channels to improve utilization. libsmctrl [24] enables per-kernel SM partitioning through TPC masking. The partitioning policies of FGPUs
5
Compute Asymmetry Analysis
5.1
SM Topology Discovery
Topology-aware partition placement requires mapping logical SM IDs to physical GPCs. We uncover the full mapping with a lightweight, two-phase probing method. 5.1.1 Phase 1: Cluster probing. Thread Block Clusters [20] are a Hopper-introduced new CUDA feature. They co-schedule a group of thread blocks from a kernel onto GPC-local SMs, enabling direct cross-block distributed shared memory access and cluster barriers without going through L2. In this case, the SMs in the same Thread Block Cluster are guaranteed to be in the same GPC. To this end, we launch kernels via cudaLaunchKernelEx with cluster sizes ranging from 3 to 8 and log each cluster’s constituent SM IDs (obtained via the %smid PTX register). SM IDs that co-occur in a cluster belong to the same GPC. With this method, we can partially uncover the GPC layout of the SMs on the GPU. 3
Xiaoze Fan et al.
Bandwidth (GB/s)
Same GPC
Different GPC
The combined method runs in under 1 minute per GPU and requires only standard GPU kernels. We validate the uncovered mapping against nvdebug on machines and confirm 100% agreement across all tested cards.
4K 2K 0
2 4 6 8 10 12 14 16 18
#SMs (a) H200
Takeaway 1: Logical SM IDs fall into two groups. Normal SMs follow an architecture-specific predefined SM-toGPC mapping, while random SMs have GPC locations determined by the chip-specific floorsweeping outcome.
2 4 6 8 101214161820
#SMs (b) B200
Figure 2. Per-GPC L2 bandwidth on H200 and B200 with varied SMs. Because each GPC in Hopper has 9 TPCs, and each GPC in B200 has 10 TPCs, the bandwidth test ends at 18 SMs on H200, and 20 SMs on B200.
5.2
Chip-specific SM Topology
5.2.1 SM Topology Variability. Figure 3-(a) shows the fully uncovered SM-to-GPC mapping on a specific H200 and B200 chip, with the blue regions indicating the normal SMs, the orange regions indicating the random SMs, and the part with red crosses indicating the floorswept SMs. Notably, Figure 3 already shows the affinity of the SMs to the NUMA partitions. Any 4 GPCs could constitute a NUMA partition. The logical SM IDs within a NUMA partition are not guaranteed to be the same across different chips. We will demonstrate the NUMA partition probing in §6. By conducting the topology discovery on multiple H200 and B200 chips of the same SKUs, we find that SM counts differ across GPCs within a chip, and GPC configurations vary across chips of the same SKU. Of the 28 H200 GPUs measured, 23 each have six 16-SM GPCs and two 18-SM GPCs. Three GPUs each have one 12-SM GPC, three 16-SM GPCs, and four 18-SM GPCs. One GPU has one 14-SM GPC, four 16-SM GPCs, and three 18-SM GPCs. The remaining GPU has one 10-SM GPC, two 16-SM GPCs, and five 18-SM GPCs. Although B200 has two additional SMs per GPC, the GPC imbalance is similar to that on H200. Additional topology results for the lower-tier H100 PCIe are provided in Appendix A. Lower-tier products from the same die floorsweep more aggressively and thus expose larger imbalance.
For instance, on H200, we can uncover the GPC layout of the SMs with IDs 0–123 (62 TPCs across 8 GPCs). These SMs reliably appear in cluster launches, while SM IDs 124–131 (4 TPCs) never appear at cluster sizes >2. We observe two groups of logical SM IDs. The SMs identified by cluster probing follow an architecture-specific predefined SM-to-GPC mapping. For convenience, we call them normal SMs. The remaining SMs have GPC locations determined by the chipspecific floorsweeping outcome. We call them random SMs. Here, “random” does not mean that these SMs are randomly scheduled at runtime. Rather, their physical GPC locations may vary across chips because of floorsweeping. 5.1.2 Phase 2: L2 bandwidth probing. To assign the remaining random SMs, we exploit the observation that SMs sharing a GPC contend on the GPC’s L2 Xbar bandwidth [25]. This means that the SMs in the same GPC have lower L2 bandwidth than the same number of SMs which are distributed to different GPCs. To get the L2 bandwidth, we use a kernel that repeatedly reads a buffer sized to exceed L1 but fit within L2, forcing all accesses to hit the L2 cache. Figure 2 shows the L2 bandwidth of SMs within the same GPC and across different GPCs on H200 and B200, respectively. As shown, the L2 bandwidth grows linearly with the number of SMs when the SMs are located in different GPCs. However, this trend does not hold for SMs within the same GPC. When the number of SMs exceeds 6 on H200 and 16 on B200, their L2 bandwidth becomes lower than that of SMs in different GPCs. With this observation, we can assign the random SMs to the GPCs with the lowest aggregated L2 bandwidth. Specifically, for each two random SMs from the same TPC with consecutive SM ids, we co-run them with the known SMs of each GPC 𝑔 to record aggregate bandwidth. If adding 𝑠 causes a bandwidth improvement smaller than a single TPC’s bandwidth, 𝑠 is contending with 𝑔’s SMs and therefore belongs to 𝑔; otherwise 𝑠 resides elsewhere. Sweeping all remaining GPCs unambiguously places every random SM.
Takeaway 2: Due to floorsweeping, SM topology varies within a chip and across chips of the same SKU, producing varied GPC imbalance. 5.2.2 Thread Block Cluster Scheduling. With the full SM topology uncovered, we re-run the cluster probing experiment on a balanced H200 chip whose 8 GPCs each contain at least 16 SMs. In this case, one GPC contains 8 normal SMs and 8 random SMs. Even on this balanced chip, the hardware never schedules blocks with cluster size larger than 2 onto the 8 random SMs. In principle, these 8 random SMs should support Thread Block Cluster scheduling with cluster sizes larger than 2. However, the driver or on-device firmware disables random SMs from Thread Block Cluster scheduling. This hides the floorsweeping induced GPC imbalance from software. Inspecting CUTLASS [26] confirms this design. Its kernel configuration logic for H200 supports a maximum of 15 thread 4
GPC
Normal TPC
Random TPC
Floorswept TPC
0
16 32 48 124 126
2
18 34 50 64 78 92 106 128
0
16 32 48 64 142
4
20 36 52 68 82 96 110 124
1
17 33 49 125 127
3
19 35 51 65 79 93 107 129
1
17 33 49 65 143
5
21 37 53 69 83 97 111 125
8
24 40 56 70 84 98 112
4
20 36 52 66 80 94 108
2
18 34 50 66 80 94 108 122
10 26 42 58 74 88 102 116 130 136
9
25 41 57 71 85 99 113
5
21 37 53 67 81 95 109
3
19 35 51 67 81 95 109 123
11 27 43 59 75 89 103 117 131 137
10 26 42 58 72 86 100 114 130
6
22 38 54 68 82 96 110
6
22 38 54 70 84 98 112 126 144
12 28 44 60 76 90 104 118 132 138
11 27 43 59 73 87 101 115 131
7
23 39 55 69 83 97 111
7
23 39 55 71 85 99 113 127 145
13 29 45 61 77 91 105 119 133 139
12 28 44 60 74 88 102 116 120
14 30 46 62 76 90 104 118 122
8
24 40 56 72 86 100 114 128 146
14 30 46 62 78 92 106 120 134 140
13 29 45 61 75 89 103 117 121
15 31 47 63 77 91 105 119 123
9
25 41 57 73 87 101 115 129 147
15 31 47 63 79 93 107 121 135 141
Partition 0
Partition 1
(a) H200
Partition 0
(b) B200
Partition 1
Figure 3. Physical SM layout across GPCs on H200 and B200. Floorsweeping disables different TPCs on each chip. The depicted left–right partition order has no physical significance, and partition 0 can correspond to either side in the actual die shot. block clusters with size 8 (14 clusters in 7 intact GPCs + 1 cluster in the smallest GPC).
Takeaway 4: To enable large Thread Block Clusters, floorsweeping forces GreenContext to allocate normal SMs first as much as possible, and only allocate random SMs at the end. This creates a priority in SM allocation.
Takeaway 3: Thread Block Cluster scheduling is constrained by chip-specific floorsweeping. Random SMs cannot be used for Thread Block Cluster with cluster size larger than 2. 5.3
5.3.2 Wasted SMs in MIG. MIG partitions a GPU into physically isolated instances. A predefined MIG profile specifies the SM count, L2 cache size, and HBM capacity of each instance. Every chip of the same SKU exposes the same set of profiles. We find that chip-specific floorsweeping causes MIG to leave some otherwise functional SMs unused. To examine this loss, we use two H200 MIG profiles, 4g.71 GB and 3g.71 GB, abbreviated as 4g and 3g. They provide 64 and 60 SMs, respectively, each with 30 MB L2 and 71 GB HBM. Their instances can coexist on one GPU, retaining the full memory capacity but leaving 8 of the GPU’s 132 SMs unused. Comparing full-GPU and MIG topologies across chips, we find that each GPC has a fixed SM composition, but the GPCs assigned to each MIG profile vary across chips. On the H200 in Figure 3, the 4g instance corresponds to the right-hand partition, containing four relatively intact GPCs. In the worst case, each of these four GPCs loses two of its 18 SMs to floorsweeping, leaving 4 × (18 − 2) = 64 SMs. The 4g profile must accommodate this worst case across chips. The depicted chip retains 68 functional SMs in this region, so MIG disables normal SMs 122 and 123 and random SMs 128 and 129 to match the 64-SM profile. The same reasoning gives 72 − 12 = 60 SMs for H200’s 3g profile if all 12 floorswept SMs fall within its 72-SM region. It also explains B200’s 4g profile, but not its 3g profile. The 80-SM region for B200’s 3g profile could theoretically lose 12 random SMs to floorsweeping, leaving 68 SMs. However, creating this instance exposes 70 SMs. One plausible explanation is that manufacturing yield permits a tighter limit of 10 floorswept random SMs in this region, guaranteeing 70 SMs. Manufacturers control floorsweeping and can impose such limits when defining a SKU.
Scheduling SMs with different mechanisms
With the uncovered SM topology, we can further analyze its impact on the different mechanisms that partition the SMs. 5.3.1 Priority in Green Context. We observe Green Contexts prioritizing normal SMs for earlier contexts, leaving random SMs for the last. Green Contexts partition SMs among kernels running concurrently on the same GPU. Each kernel is bound to a lightweight Green Context that determines its available SMs. The driver offers several SM allocation modes for Green Contexts. By default, the driver allocates SMs in GPC-aligned chunks of 8, tiled across all GPCs so that no allocation concentrates within a few GPCs. We use the default mode throughout this paper. The other modes are described in Appendix B. Since Thread Block Clusters are widely used in production kernel libraries such as CUTLASS [26], most deployments require the mode with large cluster size. This mode introduces an allocation priority across Green Contexts. Because thread block cluster with size larger than 2 cannot be scheduled onto random SMs, the driver must allocate normal SMs first to maximize the performance for current task. Therefore, all SMs assigned to earlier allocated Green Contexts are normal SMs, and the last allocated context has the random SMs. At this time, while the last Green Context may hold the same total SM count as an earlier one, only a subset of its SMs supports Thread Block Clusters larger than 2. This causes the last context to receive lower performance than expected, because random SMs cannot execute Thread Block Clusters. 5
Xiaoze Fan et al.