Conceptio › Archive › arXiv CS
arXiv CSopen access

The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations

Yikai Mao et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

arXiv:2609.34369v1 [cs.ET] 28 Sep 2026

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

Accepted to be published in: SC26 Workshops, November 15-20, 2026, Chicago, Illinois, USA

The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations Yikai Mao∗ , Yuan He∗ , Shaowen Li∗ , Masaaki Kondo∗†

∗ RIKEN Center for Computational Science, Kobe, Hyogo, Japan † Keio University, Yokohama, Kanagawa, Japan

{yikai.mao, yuan.he.uw, shaowen.li, masaaki.kondo}@riken.jp Abstract—In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirteen torus partition geometries, and six rank densities for 39-qubit simulations on 1,024 nodes) shows that partition geometry dominates runtime. All twelve algorithms run 1.73–2.31x slower on flat partitions than on near-cubic ones despite identical data transfer, proving the slowdown stems from network delivery rather than communication volume. This penalty scales with the 3D torus partition aspect ratio (runtime ∝ a0.39 , r = 0.72). Rank density is secondary, cutting runtime by 11% at 16 ranks per node only on compact geometries. Ultimately, requesting a near-cubic partition with 16 ranks per node roughly halves time-to-solution relative to flat partitions, which also consume 1.82x more energy. A simulator-free allto-all microbenchmark confirms a similar geometry penalty for collective-dominated workloads. Index Terms—quantum circuit simulation, torus networks, topology-aware placement, energy efficiency, best practices

I. I NTRODUCTION Classical simulation of quantum circuits has become a production workload on HPC systems. These simulations validate algorithms before quantum hardware execution, verify noisy device outputs against exact simulated amplitudes [1], and provide noise-free reference results for circuits that current hardware cannot yet execute reliably, making them a standing service of emerging Quantum–HPC ecosystems. At production scales, this computational workload introduces severe resource demands: simulating a 39-qubit state vector requires 8 TiB of memory, occupying on the order of a thousand nodes once communication buffers and practical per-node memory budgets are accounted for. At this scale, simulation performance is predominantly communication-bound: applying a general, non-diagonal gate to a globally distributed qubit requires exchanging state-vector data between pairs of ranks across the interconnect [2], [3]. Consequently, execution time is constrained by network performance, which depends strongly on allocation decisions made prior to job initiation. Two primary parameters govern this behavior: the userspecified number of Message Passing Interface (MPI) ranks allocated per node, and the scheduler-determined geometry of the contiguous torus partition assigned to the job. Currently, the latter parameter is rarely controlled. Eleven 1,024-node simulation jobs submitted to Fugaku with identical node

SC26 Workshops, November 15-20, 2026, Chicago, Illinois, USA 979-8-3195-1221-5/26/$31.00 ©2026 IEEE

requests received five distinct partition geometries; eight of the eleven allocations were highly elongated. Because production job packers typically optimize for machine utilization and standard submission scripts carry no information regarding application communication structure, network-bound workloads can silently lose about half their performance, and more on the flattest shapes, between otherwise identical submissions. This paper quantifies these performance variations under controlled conditions. Achieving such experimental control is non-trivial: on Fugaku, the batch scheduler accepted explicit shape requests only at partition sizes that are multiples of 96 nodes (consistent with I/O alignment) in our campaigns, and these sizes are never powers of two, whereas distributed statevector simulators require power-of-two rank counts. We resolve this constraint through a partial-use mechanism: requesting an explicitly shaped partition and executing on a powerof-two subset of the allocated nodes. Using this technique, we evaluate twelve quantum algorithms across thirteen torus geometries and six rank densities for 39-qubit simulations on 1,024 nodes, producing 764 validated configurations (936 runs minus 172 that received a partition of a different size and shape than requested). Our empirical measurements reveal a clear hierarchy between these two control variables. Partition geometry is the primary performance driver: across all algorithms, flat partitions execute 1.73–2.31× slower than near-cubic geometries. Because total injected network traffic remains constant across shapes, geometry changes how quickly the network delivers that traffic, not how much traffic there is. Rank density acts as a secondary, conditional factor: on compact partitions, allocating 16 ranks per node shortens average runtime by 11% relative to one rank per node, but it provides no measurable benefit on flat allocations. This conditional dependency means that rankdensity measurements taken without controlling placement can appear contradictory and non-monotonic. The dominant geometry effect can be modeled to first order using the partition’s aspect ratio alone, providing a practical submissiontime predictor of execution time and, because energy closely tracks time, of energy consumption. These results form an immediately deployable best practice whose reach likely extends beyond quantum simulation: a simulator-free all-toall microbenchmark shows a similar geometry penalty, so

other communication-bound workloads dominated by all-toall or transpose-style collectives on contiguous torus partitions should benefit as well. Distributed quantum circuit simulation is a demanding and timely instance: a single 39-qubit job moves 0.3 to 3.2 PB across the interconnect, depending on the circuit. Our primary contributions are as follows: • To our knowledge, the first controlled characterization of partition-geometry effects on distributed quantum circuit simulation on a production supercomputer, including the allocation methodology that makes such controlled experiments feasible (Sections III and IV). • Empirical identification of the placement hierarchy: partition geometry serves as a ∼2× performance lever consistent across all twelve algorithms, whereas rank density acts as a conditional lever (11% shorter runtime) dependent on compact geometry (Sections III-B, IV-B, and IV-C). • Formulation of a single-parameter surrogate model (runtime ∝ α0.39 , r = 0.72) that gives workload managers an evaluable submission-time cost model, and of an associated best-practice execution strategy that improves time-to-solution by roughly 2× relative to flat partitions, which also consume 1.82× more energy (Sections III-C, III-D, IV-D, and IV-E). The paper is structured as follows: Section II reviews related work. Section III describes the placement hierarchy, recommended practices, and allocation mechanism. Section IV presents our empirical evaluation. Section V discusses limitations and system implications, and Section VI concludes. II. BACKGROUND AND R ELATED W ORK A. Distributed State-Vector Simulation and Its Communication State-vector simulation stores all 2n complex amplitudes of an n-qubit register and applies gates as sparse linear operators over them [4]–[6]. Beyond roughly 30 qubits (16 GiB) on a 32 GiB Fugaku node, the state vector exceeds a single node’s memory and must be distributed across multiple nodes. The qubits are consequently partitioned into rank-local and global qubits; a non-diagonal gate on a global qubit pairs each rank with the partner whose index differs in that qubit’s bit, and the two exchange state-vector data, over the interconnect whenever the partners reside on different nodes [2], [3]. Large-scale state-vector simulations on petascale systems [2], [7] and multi-GPU simulators built on the cuQuantum SDK [8] report that performance at scale is limited largely by data movement rather than arithmetic, whereas tensor-network simulators [1], [9] avoid much of this communication at the cost of other trade-offs. An extensive body of literature addresses this bottleneck from the application layer by reducing data movement or memory footprint via gate scheduling, cache blocking, and gate fusion [10], qubit reordering [11], data compression [12], and communication-avoiding partitioning [13]. Recent benchmarks on GPU clusters also identify interconnect performance as a key scaling limit [14]. While prior simulator work focuses

on optimizing the traffic generated by the application, it leaves open the question of how efficiently the underlying network delivers that traffic. On a torus interconnect that gives each job a contiguous, private partition, this delivery efficiency depends strongly on the partition’s geometry, which is fixed before computation begins. B. Topology-Aware Placement in HPC Systems The impact of job placement on performance is well established for classical workloads across various network topologies. On Cray systems, Bhatele et al. observed runto-run performance variations exceeding 30%, caused mainly by network contention from neighboring jobs, whereas Blue Gene systems, which give each job an isolated torus partition, showed little variability [15]. Consequently, scheduler research has emphasized topology-aware allocation algorithms that minimize fragmentation while giving communication-sensitive jobs compact, convex shapes [16]. This placement sensitivity is also observed on other interconnects: job-interference studies on dragonfly networks trade locality against hotspot avoidance [17], other studies link performance predictability to allocation isolation [18], and measurements on a modern dragonfly system show that hardware congestion control can largely mitigate such interference [19]. What distinguishes systems that allocate contiguous, private torus partitions is that the placement variable is explicit and low-dimensional. On Fugaku, each large job receives a contiguous block of the Tofu Interconnect D network, which it sees through Tofu’s virtual three-dimensional torus rank mapping [20]. While this contiguous allocation largely eliminates interference from neighboring jobs, it leaves the partition’s own geometry as a dominant, yet largely unexamined, performance factor. Simulation studies of fat-tree clusters found that the cost of contiguous allocation is compensated when job runtimes improve by roughly 20–30% [21], and work on IBM Blue Gene/Q showed that relaxing network allocation constraints improves scheduling, especially when it accounts for each job’s communication sensitivity [22]. Oltchik and Schwartz further showed that the shape of an allocated partition can cause avoidable contention and derived a predictor from partition structure [23]; our measurements complement theirs with quantum-simulation workloads on Fugaku’s Tofu D network. This topology sensitivity has re-emerged prominently in modern accelerated architectures: for example, Slurm’s Block Topology plugin for NVIDIA GB200 NVL72 systems enforces contiguous placement to restrict all-to-all communication to high-bandwidth NVLink domains rather than crossing standard network fabrics [24]. Apart from these few studies, prior topology-aware research has primarily optimized rank-to-node mappings within a fixed allocation or evaluated allocation heuristics via simulation. Controlled, applicationlevel measurements of allocation geometry on production systems remain scarce, particularly for distributed quantum circuit simulation, where every non-diagonal global-qubit gate exchanges large state-vector slices between pairs of ranks that

short diameter long diameter, narrow bisection 8

12

2 9 16 near-cubic 8×9×16 aspect ratio α = 2

48 flat 2×12×48 aspect ratio α = 24

1024 nodes used (power of two)

128 idle (padding)

same 1152-node allocation, same injected bytes:

flat runs ∼2× slower

Fig. 1. Two possible geometries for a 1,152-node Tofu allocation, drawn roughly to scale, with thin axes widened and the long axis shortened for legibility. The simulation runs on a power-of-two subset of 1,024 nodes (blue); 128 nodes idle as padding (gray), the same 12.5% overhead that the scheduler’s automatic padding imposes on unshaped requests. Both partitions carry identical traffic, but the flat one drains it through a longer diameter and a narrower bisection. Averaged over all flat geometries (α ≥ 16), runs take about twice as long as on near-cubic ones (the ∼2× in the figure); the drawn pair differs by 3.2×.

may lie far apart in the partition, a severe stress test for network placement. C. Benchmarking and Workload Management Standardized benchmark suites and circuit generators [25]– [27] define what to run when evaluating quantum circuit simulation, and published evaluations report scaling with qubit count, node count, and rank configuration. What common benchmarking practice rarely controls, or even records, is the geometry of the allocation each measurement ran on. Our controlled grid shows that geometry alone changes runtime by about 2× at identical traffic (Section IV-B) and determines whether tuning rank density pays off (Section IV-C), so placement belongs in the measurement protocol of any communication-bound benchmark, alongside the variability controls motivated in the placement literature above. Systems that schedule quantum workloads alongside classical ones [28]–[31] focus on sharing quantum resources and generally leave the network placement of classical jobs to the underlying batch scheduler, which yields default allocations like those in Section III-A. The aspect-ratio surrogate of Section III-C addresses this gap with an empirical, singleparameter cost model evaluable at job submission, and the practice of Section III-D makes the benefit available to users without any scheduler changes. III. F ROM P LACEMENT L OTTERY TO B EST P RACTICE A. The Placement Lottery of Default Allocation When the node count is set by memory, that is, by the state vector and its communication buffers rather than by computation, distributed state-vector simulation is fundamentally limited by communication (Section II). At 39 qubits, the 8 TiB state vector is distributed over 1,024 Fugaku nodes (8 GiB per node), leaving room for communication buffers. Under these operating conditions, execution time is governed by two parameters fixed before execution begins: the geometry of the

TABLE I G EOMETRIES ASSIGNED BY DEFAULT PLACEMENT ACROSS ELEVEN 1,024- NODE SUBMISSIONS WITH IDENTICAL NODE REQUESTS ON F UGAKU ; THE SCHEDULER PADDED EACH REQUEST TO THE 1,152- NODE PARTITIONS LISTED . T HE ASPECT RATIO α IS THE LONGEST AXIS DIVIDED BY THE SHORTEST. Assigned geometry

Aspect ratio α

Frequency

2×12×48 2×18×32 4×9×32 8×6×24 8×9×16

24 (flattest) 16 8 4 2 (near-cubic)

5 3 1 1 1

contiguous torus partition assigned to the job (Figure 1), and the number of MPI ranks allocated per node. Currently, users rarely specify or control partition geometry. We characterize geometry using the partition aspect ratio, defined as α = max(A, B, C)/ min(A, B, C) for a contiguous sub-torus of dimensions A × B × C, where α = 1 denotes a perfect cube and larger values indicate increasingly elongated allocations. Table I reports the geometries assigned to eleven 1,024-node submissions (different circuits and rank densities) where only the total node count was requested. The scheduler padded each request to a 1,152-node contiguous partition and assigned five different geometries, presumably depending on free-space availability across the torus. Flat allocations (α ≥ 16) account for eight of the eleven submissions, with the flattest geometry (α = 24) assigned five times. This pattern is consistent with packing policies that preserve compact free regions while filling elongated remnants. Consequently, default placement is neither reproducible nor performanceaware. Because flat partitions (α ≥ 16) execute this workload roughly twice as slowly as near-cubic geometries (as shown in Section IV), standard job submissions can forfeit about half of their potential performance, and more on the flattest geometries. Moreover, standard submission interfaces do not capture an application’s communication characteristics, so schedulers cannot optimize for this behavior automatically. The remainder of this section analyzes data from 764 controlled configurations (Section IV) to establish a twolevel placement hierarchy (Section III-B), propose a scalar performance predictor (Section III-C), outline recommended user practices (Section III-D), and detail the allocation mechanism that enables controlled geometry experiments on an unmodified production scheduler (Section III-E). B. Geometry Dominates and Constrains Rank Density Partition geometry is the primary performance determinant. Across all twelve algorithms evaluated, flat partitions (α ≥ 16) increase execution time by 1.73–2.31× (mean 2.04×) compared to near-cubic partitions (α ≤ 3, termed compact or cubic below), consistently across gate counts ranging from 91 to 1,404. This slowdown does not come from extra traffic: for a given algorithm and rank density, the injected data volume is the same on every geometry (Section IV-B), because the exchange pattern depends on the amplitude distribution, not

on the physical mapping of ranks. Partition geometry impacts performance by modifying the partition diameter and bisection bandwidth, changing how quickly the network drains those bytes, as Figure 1 illustrates. Because the bottleneck lies in the network rather than in any circuit-specific computation, all twelve algorithms suffer nearly the same slowdown. Rank density acts as a secondary factor whose effectiveness is conditional on geometry. On near-cubic partitions, allocating 16 ranks per node (3 OpenMP threads per rank) yields an average 11% shorter runtime than one rank per node. Conversely, on flat partitions, rank-density adjustments yield no measurable benefit (within 2%, Table IV), consistent with execution being bounded by inter-node network delivery, which changes in rank density cannot alleviate. Oversubscription (32 ranks × 2 threads = 64 threads on 48 cores) performs worse on average than 16 ranks per node, although two algorithms (QAOA, QW) achieve optimal runtimes under this configuration (Table III); as a general default, oversubscription should be avoided. This conditional relationship can make rank-density tuning appear inconsistent wherever placement goes uncontrolled, as in our own earlier campaign (Figure 3), where node count also varied: under default scheduler policies, partition geometry varies arbitrarily between runs. Consequently, measurements across different submissions conflate evaluations on compact geometries (where density impacts performance) with evaluations on flat geometries (where it does not). Our controlled grid isolates these variables. C. Aspect Ratio as a Predictive Surrogate Among candidate geometric metrics, the partition aspect ratio, introduced in Section III-A, is the strongest predictor of execution time. After normalizing scale differences across algorithms and densities, the correlation between log runtime and log α across all 764 configurations is r = 0.72, outperforming inverse bisection width (the reciprocal of the product of the two shorter axes) (r = 0.67) and inverse shortest axis length (r = 0.61). A power-law regression across the dataset yields: T (α) ≈ T1 · α 0.39 , (1) where T1 is the runtime extrapolated to α = 1 (a perfect cube, not realizable at this node count) for a given algorithm and density. Eq. (1) predicts a 2.6× execution-time penalty when shifting from α = 2 to α = 24 and a 2.2× ratio between the near-cubic and flat classes, somewhat above the observed 2.04×; individual geometries depart from the fit (Section IV-D). Partition elongation predicts performance degradation somewhat better than minimum-cut bandwidth alone. Because these metrics are strongly correlated across the 13 geometries (r = 0.87–0.97), and a mean-hop-distance proxy (the sum of the axes) scores similarly (r = 0.70), the data cannot separate path length from bisection bandwidth. Application-level metrics carry comparatively little predictive weight (gate count vs. topology sensitivity: r = −0.27), allowing this surrogate model to function without per-application calibration beyond the scaling constant T1 .

Because this surrogate metric can be evaluated at job submission using only partition dimensions, it provides a simple cost model for distributed state-vector simulation on Fugaku. While an empirical correlation of r = 0.72 indicates moderate scatter around the regression line, including geometry-specific departures (discussed in Section IV-D), Eq. (1) serves as a lightweight, submission-time screening heuristic rather than an exact analytical model. Given a set of available torus regions, a scheduler could score feasible placements by α0.39 , weighing the predicted slowdown of an elongated allocation against the wait time for a compact region to co-optimize time-to-solution and energy consumption. Flat partitions also consume 1.82× more energy than near-cubic ones (Section IV-E), so compact placements improve both performance and energy efficiency. While building an autonomous allocator is outside the scope of this paper, our measurements provide the empirical cost model needed to support one. D. Recommended Execution Practice For end users, these empirical findings suggest a straightforward operational procedure: 1) Request geometry explicitly: Specify an I/O-aligned, near-cubic partition (α ≤ 3; in our study, 8×9×16 or 8×18×8, as 6×12×16 ran markedly slower) and execute using power-of-two partial node utilization (Section III-E). 2) Optimize rank density: Allocate 16 ranks per node as the standard baseline; tune density only when both the workload and the geometry are known (Section IV-C). 3) Handle flat allocations: If only flat partitions are available, rank-density adjustments will not improve performance. Users should either defer the job until a compact region is available (using Eq. (1) to quantify the cost of immediate execution) or run immediately at any density (e.g., R = 16). Adopting these practices improves time-to-solution by roughly 2× relative to flat partitions, which default placement assigned to eight of eleven submissions and which also consume 1.82× more energy than near-cubic ones (Section IV). This approach requires no modifications to the simulation software or underlying scheduler and can be applied today on Fugaku and on other systems whose schedulers accept explicit partition shapes, provided the allocated shape is checked (Section III-E). E. Obtaining Defined Geometries on Production Schedulers Conducting controlled geometry experiments on a production system requires overcoming a structural scheduling constraint. On Fugaku, a standard 1,024-node submission is padded automatically to a larger contiguous partition of scheduler-determined geometry, 1,152 nodes in all eleven default submissions of our campaign (occasionally other sizes, such as a 1,200-node case observed in the microbenchmark campaign of Section IV-F, Table V). Requests for explicit shapes at unaligned counts (e.g., 16×8×8 = 1,024) are overridden by scheduler padding rules. In our campaigns, explicit

geometries were honored only when the requested node count was a multiple of 96 (1,152 = 12 × 96), consistent with the system’s I/O-alignment granularity. As a result, alignable counts such as 1,152 = 27 32 are never powers of two. However, distributed state-vector simulators require a powerof-two rank count to partition 2n amplitudes evenly. The two constraints therefore have no common solution: no aligned node count, fully used at a uniform number of ranks per node, yields a power-of-two rank count, and neither constraint can be relaxed. This incompatibility helps explain why controlled placement studies of this workload have been impractical on Fugaku to date. We resolve this incompatibility through partial use: requesting an explicitly shaped, I/O-aligned partition (1,152 nodes) and executing the simulation on a power-of-two subset (1,024 · R ranks across 1,024 nodes at R ranks per node), leaving the remaining 128 nodes idle. This allows execution within a controlled, user-specified geometry at the cost of a 12.5% idle-node allocation overhead. Crucially, this overhead matches the padding cost that default placement imposed on all eleven unshaped submissions of Table I; partial use thus converts unmanaged padding into an experimental control variable. Even against a hypothetical zero-padding 1,024node allocation that lands on a flat partition, the 2.04× speedup outweighs the idle nodes, yielding a net advantage of (1,024/1,152) × 2.04 ≈ 1.8× in completed simulation work per allocated node-hour. We requested thirteen distinct 1,152node geometries, each of which the scheduler granted in at least some runs, spanning aspect ratios from α = 2 (nearcubic 8×9×16) to α = 24 (flat 2×12×48); these provide the experimental basis for Section IV. The scheduler did not always honor these requests, however: 172 of our 936 runs (18%) received a partition of a different size and shape and were excluded from all analyses, and 80 runs received the requested dimensions in a different axis order and were retained as separate orientation variants (e.g., 16×9×8 for a requested 8×9×16). IV. E VALUATION This section details the experimental methodology supporting Section III and presents the full evidence: the setup (IV-A), the geometry effect (IV-B), the shape–density interaction (IVC), the predictor comparison (IV-D), energy scaling (IV-E), and a simulator-free microbenchmark (IV-F). A. Experimental Setup Platform. Experiments were conducted on Fugaku, which features 48-core Fujitsu A64FX processors with 32 GiB of HBM2 memory (1,024 GB/s bandwidth), connected by the Tofu Interconnect D 6D mesh/torus network [20] providing ten 6.8 GB/s links per node. Evaluations used the MPI-parallel implementation of Qulacs for A64FX [3]. Hardware counters provided measurements for node power consumption (to calculate total energy) and Tofu network user traffic (to measure communication volume).

TABLE II W ORKLOAD : TWELVE QUANTUM ALGORITHMS , ALL AT n = 39 QUBITS (8 T I B STATE , 8 G I B PER NODE ON 1,024 NODES ), SORTED BY CATEGORY AND ALGORITHM NAME . C ATEGORIES FOLLOW Q-G EN [25]. A LL ENTRIES ARE FIXED - DEPTH REPRESENTATIVE KERNELS ; E . G ., VQE IS A SINGLE ANSATZ EVALUATION , AND THE S HOR - STYLE KERNEL USES CNOT PATTERNS AND A TRUNCATED QFT IN PLACE OF MODULAR - EXPONENTIATION ARITHMETIC .

Category

Algorithm

Gates

Communication

Quantum Key Distribution (QKD) Quantum Teleportation (QT)

91 114

Fourier

Quantum Phase Estimation (QPE) Shor-style kernel (Shor)

154 356

Query

Bernstein–Vazirani (BV) Deutsch–Jozsa (DJ)

96 115

Search

Grover’s Algorithm (Grover) Quantum Counting (QC) Quantum Walk (QW)

735 192 770

Variational

Quantum Approx. Optim. (QAOA) Var. Classifier (VQC) Var. Quantum Eigensolver (VQE)

507 1,404 462

Workload. To avoid confounding geometric sensitivity with problem size, every configuration simulated n = 39 qubits. The resulting 8 TiB state vector allocates 8 GiB per node across 1,024 nodes, a node count set by memory, as in production distributed quantum simulation. The evaluation suite comprises twelve algorithms spanning the query, communication, variational, Fourier, and search categories of the Q-Gen generator [25] (Table II). Because standard benchmarking suites do not provide these circuits at 39 qubits, each circuit was generated programmatically at full width as a representative fixed-depth kernel of each algorithm rather than a complete textbook implementation (e.g., QPE omits the final inverse QFT, and the quantum-counting kernel is a phase-estimation circuit). For high-depth algorithms (Grover, QC, QPE, QW, Shor), iteration counts were fixed at small constants to ensure jobs completed within scheduler wall-time limits. Gate counts range from 91 (QKD) to 1,404 (VQC), with every circuit operating on the full 39-qubit register. At 39 qubits, execution time is largely determined by global state-vector exchanges over the network. Because production runs of iterative algorithms (e.g., repeated VQE or QAOA ansatz evaluations) repeat these same exchanges, the geometry penalty measured on one kernel pass is expected to carry over to end-to-end execution time roughly proportionally. For example, the mean VQC pass drops from 1,576 s on flat partitions to 912 s on near-cubic ones (Table III), saving about 213 node-hours per pass on a 1,152-node allocation; an iterative run repeats this saving at every pass, and obtaining it requires only an explicit, verified shape in the job request (Section III-E). Sweep design. The factorial experimental grid comprises 12 algorithms × 13 geometries × 6 rank densities = 936 configurations. Each configuration was executed as an independent batch job using the explicit placement mechanism described

B. Partition Geometry Dominates Performance Table III quantifies the primary level of the placement hierarchy. For every evaluated algorithm, the fastest execution time across its retained (geometry, density) combinations occurred on a near-cubic geometry: the Best cell column contains only near-cubic partitions: 8×9×16 or its rotated orientation 16×9×8 (eight algorithms, typically at R = 16) and 8×18×8 at R = 1 (the remaining four). The ratio of mean execution time on flat partitions to near-cubic partitions falls within a band of 1.73–2.31× (mean 2.04×), despite best-configuration runtimes spanning 52 s (QT) to 709 s (QC) and gate counts ranging from 91 to 1,404. Figure 2 illustrates this relationship continuously across aspect ratios. When aggregated across all algorithms, normalized runtime trends upward with aspect ratio, following the fitted power law of Section III-C with substantial per-geometry scatter (r = 0.72). These differences arise from how the network delivers the traffic, not from its volume. For a given algorithm, total user data injected into the Tofu network remains invariant across partition geometries: per-job interconnect counters confirm that, at 16 ranks per node, DJ injects 668.5 TB and VQC injects 3,166.7 TB, remaining constant across all retained geometries (13 for DJ, 12 for VQC) to within 10−6 %. Because communication volume is fixed by the amplitude distribution, elongated geometries must transfer identical byte volumes across longer diameters and narrower bisections, requiring roughly twice the execution time.

TABLE III P ER - ALGORITHM RESULTS , SORTED BY ALGORITHM NAME , OVER RETAINED RUNS . B EST CELL OVER ALL RETAINED ( GEOMETRY, RANK DENSITY R) CONFIGURATIONS ; CLASS MEANS AVERAGE OVER ALL RETAINED ( GEOMETRY, R) CELLS OF THE CLASS ( CUBIC : α ≤ 3; FLAT: α ≥ 16; 6–18 CELLS PER CLASS AND ALGORITHM ); FLAT / CUBIC RATIOS FOR EXECUTION TIME AND ENERGY TO SOLUTION . R ATIOS ARE COMPUTED FROM UNROUNDED MEANS . Best cell

Mean (s)

Flat/cubic

Algorithm

(s) geometry, R

Cubic

Flat

Time Energy

BV DJ Grover QAOA QC QKD QPE QT QW Shor VQC VQE

128 193 413 236 709 59 497 52 300 208 660 293

8×18×8, 1 8×18×8, 1 8×9×16, 16 16×9×8, 32 8×18×8, 1 16×9×8, 16 8×18×8, 1 16×9×8, 16 8×9×16, 32 8×9×16, 16 8×9×16, 16 16×9×8, 16

216 357 571 333 1,411 85 1,044 70 404 284 912 430

482 750 1,221 710 2,911 184 2,004 131 935 516 1,576 840

2.23× 2.10× 2.14× 2.13× 2.06× 2.16× 1.92× 1.88× 2.31× 1.82× 1.73× 1.95×

1.95× 1.95× 1.92× 1.93× 1.90× 1.69× 1.82× 1.57× 2.05× 1.69× 1.62× 1.73×

2.04×

1.82×

Mean

runtime / cubic-class mean (log)

in Section III-E, ensuring that geometry and rank density remained unconfounded and that hardware counters reflected individual configuration performance. Evaluated rank densities included R ∈ {1, 2, 4, 8, 16, 32}, using 48/R OpenMP threads per rank for R ≤ 16 (close binding, all 48 cores occupied) and 2 threads per rank at R = 32, representing the sole oversubscribed configuration (64 threads on 48 cores). The power-oftwo rank constraint excluded R = 24 and R = 48, and the scheduler limit of 48 processes per node excluded R = 64. MPI ranks used the scheduler’s default rank-to-node mapping in every cell, with no explicit rank-map files supplied, so the mapping policy was uniform across the grid. During execution, each job constructed the circuit, performed one warm-up pass, and recorded the execution time of one full-circuit update pass; across the grid, the warm-up pass agreed with the timed pass to a median difference of 0.6%. A run counted as completed only if its timed pass finished and its result record confirmed n = 39 and exactly 1,024 distinct execution nodes. Failed or timed-out jobs were resubmitted until every cell had a completed run. For each cell, the latest completed run was retained only if the scheduler record showed that its allocated partition matched the requested dimensions, in any axis order; 172 of the 936 cells (18%) failed this check and were excluded, leaving 764 runs on which all application-grid results are based. Alongside the application grid, a pure MPI all-to-all microbenchmark provided a simulator-free evaluation of geometry effects (setup and results in Section IV-F).

3.5

Algorithms

3.0

QKD QT QPE Shor

2.5 2.0

BV DJ Grover QC

QW QAOA VQC VQE

1.5

1.0 0.8

Geometry 8×9×16 8×18×8 6×12×16 8×6×24

0.6

2

4

8 partition aspect ratio α (log)

rotated 4×18×16 6×6×32 4×12×24 4×9×32

2×24×24 4×6×48 2×18×32 8×3×48 2×12×48

16

24

Fig. 2. Runtime versus partition aspect ratio, the 173 (algorithm, allocated 1 partition) pairs with retained runs (density-averaged, normalized to each algorithm’s mean over near-cubic cells, α ≤ 3), with the fitted power law T ∝ α0.39 . Color identifies the algorithm (one color family per category of Table II) and marker shape the geometry; hollow markers show runs that received the requested dimensions in a different axis order. The 13 geometries yield 11 distinct aspect ratios because two pairs coincide: α = 12 (2×24×24 and 4×6×48) and α = 16 (2×18×32 and 8×3×48). Most algorithms follow the upward trend, with geometry-specific departures discussed in Section IV-D.

We can now assess the performance impact of default placement (Table I). Default placement assigned flat geometries (α ≥ 16) to eight of the eleven submissions, which run 2.04× slower than near-cubic ones on average. A default submission therefore typically takes roughly twice as long as the recommended execution practice (Section III-D), a penalty that rank-density tuning alone cannot overcome.

TABLE IV I NTERACTION OF SHAPE CLASS AND RANK DENSITY R: RUNTIMES NORMALIZED TO EACH ( ALGORITHM , GEOMETRY ) PAIR ’ S OWN MEAN OVER ITS RETAINED R ( LOWER IS BETTER ), AVERAGED OVER THE ALGORITHMS AND GEOMETRIES OF EACH CLASS . B OLD MARKS THE BEST CELL IN THE TABLE .

placement, node count, and resource group varied with density. Under that protocol, runtime swings by more than twofold and the apparent density trend reverses; controlling geometry and node count isolates the 11% effect on compact partitions. D. Predictive Power of Aspect Ratio

Shape class

R=1

2

4

8

16

32

Near-cubic (α ≤ 3) Flat (α ≥ 16)

1.06 0.99

1.03 1.00

1.01 0.99

0.99 1.00

0.95 1.01

0.97 1.01

uncontrolled (8,192 ranks, nodes vary)

1.6

runtime / per-curve mean

controlled, near-cubic (α ≤ 3) controlled, flat (α ≥ 16)

1.4

1.2

1.0

0.8

0.6 1

2

4

8

16

32

ranks per node R

Fig. 3. Rank-density curves under the controlled protocol (near-cubic vs. flat 1 classes) and under the conventional uncontrolled protocol of our earlier 15algorithm campaign (total ranks fixed at 8,192, so node count and placement vary with density); all three curves are normalized to their own means. The uncontrolled curve swings by 2.15× and reverses direction, but it confounds placement with node count; the controlled curves isolate the conditional effect.

C. Interaction Between Shape and Rank Density Table IV quantifies the secondary level of the placement hierarchy. Because each geometry’s retained density runtimes are expressed relative to their own internal mean, the geometry effect is divided out, so the residual variation reflects rank density; restricting to the (algorithm, geometry) pairs with all six densities retained changes no entry by more than 0.02. On near-cubic geometries, rank density has a clear effect: shifting from R = 1 to R = 16 reduces mean execution time by 11% (1.06 vs. the bolded 0.95), while the oversubscribed R = 32 is slower than R = 16. Conversely, on flat geometries, the performance curve remains essentially level (within 2%). This asymmetry is the conditional behavior summarized in Section III-B; a plausible explanation is that rank density mainly changes intra-node execution, whose gains are visible on compact partitions but masked by longer network delivery times on flat ones. Aggregated means also mask cell-level heterogeneity: on 8×18×8, the fastest retained cell for BV, DJ, QC, and QPE is at R = 1 (Table III), whereas on 8×9×16 R = 16 beats R = 1 for every algorithm with both runs. Therefore, R = 16 serves as the best general default, and density is worth tuning only when both the workload and the target geometry are known. Figure 3 sets these controlled curves against our earlier uncontrolled campaign, in which

After normalizing scale across algorithms and densities, the correlation between log runtime and log α across all 764 cells is r = 0.72, compared to r = 0.67 for inverse bisection width and r = 0.61 for inverse shortest axis length. The fitted exponent of 0.39 (Eq. (1)) predicts a 2.6× spread between the α = 2 and α = 24 endpoints and a 2.2× ratio between the near-cubic and flat classes, against an observed 2.04×; the measured endpoint pair differs by 3.2×, reflecting the geometry-specific departures discussed below. Circuit-level features fail to predict topology sensitivity reliably: correlation with gate count is weak and negative (r = −0.27), and individual gate-heavy workloads can still show high shape sensitivity (e.g., QW at 770 gates shows a 2.31× ratio), because gate count does not reflect the proportion of globalqubit communication. The power law captures the dominant trend, but individual geometries depart from it (Figure 2). Averaged over algorithms, 6×12×16 (α ≈ 2.67) runs 33% slower than Eq. (1) predicts (range 5–79%), and 6×6×32 (α ≈ 5.33) runs 61% slower (17–148%). Conversely, 8×3×48 runs 36% faster than predicted (averaged over the eight algorithms with retained runs), and faster than 2×18×32 at the same α = 16. Orientation matters too: the same dimensions in a different axis order ran up to about 50% faster or slower at equal rank density (e.g., 6×24×8 vs. 8×6×24), although α is identical. Partitions with equal aspect ratio can therefore perform differently, so α alone does not capture all geometric effects. Pinning down the cause would require per-link traffic counters, which are not exposed at this scale (Section V-A). We leave to future work a multivariate model that combines diameter, bisection width, and the mapping of the logical partition onto Tofu’s sixdimensional physical coordinates [20], validated against linklevel measurements. In its current form, the scalar α already serves as a submission-time screening heuristic that requires no routing knowledge. E. Energy Scaling Characteristics Table III and Figure 4 compare flat-to-cubic performance ratios for time and energy consumption. Energy penalties (1.57–2.05×, mean 1.82×) consistently remain below execution time penalties (mean 2.04×). On flat partitions, processors spend more time waiting for network communication, which is consistent with a roughly 6% lower average node power (job energy divided by job elapsed time) than on near-cubic allocations (power ratio 0.92–0.98 across algorithms). Total energy calculations integrate all allocated nodes, including idle padding, identically across both classes. Energy scales sublinearly with time because of the lower node power on flat partitions and because fixed per-job phases outside the timed pass add similar energy in both classes. Even so, near-cubic

all-to-all bandwidth per rank (GiB/s, log)

2.4

energy = time one point per algorithm

energy penalty, flat/cubic

2.3 2.2 2.1 2.0 1.9 1.8 1.7 1.6

9 allocations fit BW ∝ α−0.44

2.5 2.0

1.5

1.0 0.8

0.6 2

1.5 1.5

1.6

1.7

1.8

1.9

2.0

2.1

2.2

2.3

4

8

16

24

allocated partition aspect ratio α (log)

2.4

time penalty, flat/cubic

Fig. 4. Energy penalty versus time penalty of flat relative to near-cubic 1 partitions, one point per algorithm. Every algorithm falls below the diagonal: energy grows sub-linearly with time. TABLE V A LL - TO - ALL MICROBENCHMARK : REQUESTED VERSUS ACTUALLY ALLOCATED GEOMETRY ( RECOVERED FROM SCHEDULER RECORDS ) AND SUSTAINED PER - RANK BANDWIDTH AT 1 M I B MESSAGES ( ONE RANK PER NODE ), SORTED BY ALLOCATED ASPECT RATIO . D EFAULT SUBMISSIONS REQUEST NO SHAPE . N ODES : ALLOCATED ; U SED : NODES ( RANKS ) THAT RAN THE BENCHMARK .

Requested

Allocated

Nodes

Used

α

GiB/s

16×16×8 (default) 16×8×8 32×16×4 32×16×2 32×16×1 32×32×1 32×32×2 (default)

16×9×16 10×15×8 16×9×8 4×18×32 2×18×32 2×18×32 2×33×32 2×33×32 2×24×48

2,304 1,200 1,152 2,304 1,152 1,152 2,112 2,112 2,304

2,048 1,024 1,024 2,048 1,024 512 1,024 2,048 2,048

1.8 1.9 2.0 8.0 16.0 16.0 16.5 16.5 24.0

1.72 2.60 1.97 0.77 0.79 0.77 0.86 0.86 0.63

partitions improve both metrics, reducing execution time by roughly half and energy consumption by approximately 45% compared to flat partitions, which default placement assigned most often in our campaign (Table I). F. Microbenchmark Corroboration A standalone MPI all-to-all benchmark shows that the interconnect alone reproduces this geometric sensitivity, without any simulator software. We ran it (1 MiB messages, 20 timed iterations after warm-up, one rank per node) across allocations whose actual geometries were recovered from scheduler records; requested and allocated geometries frequently differ because the scheduler’s padding rules override shape requests at unaligned node counts (Section III-E; Table V; e.g., a 32×16×4 request was placed as 4×18×32). Across nine allocations (seven distinct geometries) of 1,152 to 2,304 nodes, of which 512 to 2,048 ran the benchmark, near-cubic geometries (α ≈ 2) sustained 1.7–2.6 GiB/s per rank, whereas flat allocations (α = 16–24) sustained 0.63–0.86 GiB/s, a 2.7× class difference; the slowest cell (0.63 GiB/s) is an unshaped

Fig. 5. All-to-all per-rank bandwidth versus actually allocated aspect ratio 1 for the nine allocations of Table V, with the fitted power law BW ∝ α−0.44 (r = −0.94). Two allocations coincide at α = 16.5.

default request that the scheduler placed at 2×24×48. Across all nine allocations, bandwidth falls off as a power law in aspect ratio with exponent −0.44 (r = −0.94, Figure 5). This exponent is close to the application-level 0.39 in Eq. (1); the small difference is within the uncertainty of a nine-point fit and is in the direction expected if applications dilute network penalties with local computation. Although the number of participating nodes varies by 4×, making the exponent indicative rather than exact, this simulator-free measurement reproduces a similar systematic power-law dependence on the same hardware platform, supporting the network deliveryspeed mechanism. V. D ISCUSSION AND L IMITATIONS A. Limitations Five experimental limitations qualify these claims. (1) Using a 1,024-node subset within a 1,152-node prism creates a slightly irregular communication geometry compared to the nominal shape; however, the microbenchmark of Section IV-F reproduces the same systematic trend across subsets that use 44–97% of their allocations, which suggests that this irregularity does not drive the effect. (2) Each configuration cell is timed using a single full-circuit pass within one job allocation. The warm-up pass agrees with the timed pass to a median difference of 0.6%, with 25 of the 764 cells differing by more than 5% (at most 13%); we did not repeat allocations, so run-to-run variability across allocations is not quantified. (3) All circuits are fixed-depth representative kernels: deep algorithms use truncated iteration counts, and the Shor-style kernel omits modular exponentiation; while these kernels exercise the same global state-exchange pattern that production runs repeat (Section IV-A), they do not reflect full end-to-end algorithmic runtimes lasting hours. (4) Link-level hardware counters are not exposed at this scale, so network-bound behavior is inferred from invariant traffic volume, geometrydependent execution times, and the microbenchmark results of Section IV-F, rather than direct link-level measurements.

(5) The recommended practice trades queue wait time for execution time: an explicit shape request may wait longer in job queues than a default submission. In our campaign, compact shape requests showed no such penalty relative to flat ones: among our 1,152-node requests, near-cubic jobs often started ahead of earlier-submitted flat jobs, whereas flat jobs only rarely started ahead of earlier-submitted near-cubic ones. A broader comparison against unshaped submissions and across system load conditions remains future work. B. Generalization Beyond Fugaku While the numerical constants reported here are specific to Fugaku, the underlying system-level principles generalize. Any architecture allocating contiguous partitions on a torus network presents a similar low-dimensional placement variable. Workloads with placement-invariant communication volumes and communication-bound execution should exhibit systematic geometry penalties governed by system-specific power-law exponents. The methodology of Sections III-E and IV-A, including partial node utilization to resolve alignment constraints, transfers directly to similar environments. On dragonfly or fat-tree interconnects, allocations are typically non-contiguous and share network links with other jobs, so the placement variable is neither low-dimensional nor isolated from inter-job contention (Section II); on those systems, an equivalent surrogate model would score allocations using measured congestion states rather than static geometry. Nor is the effect specific to quantum simulation: the microbenchmark of Section IV-F reproduces the geometry penalty with a plain MPI all-to-all and no simulator code, so workloads dominated by global all-to-all or transpose-style collectives, such as parallel three-dimensional FFTs and spectral solvers, are likely to benefit from the same shape request; whether the conditional rank-density effect carries over remains to be tested. Within quantum simulation itself, because distributed state-vector simulators generally rely on similar pairwise stateexchange patterns, the geometry penalty should also appear for GPU-based simulators on torus-like interconnects, even if specific power-law exponents differ. Furthermore, the emergence of multi-tier accelerated systems, such as NVIDIA GB200 clusters using the Slurm Block Topology plugin [24], underscores that topology-aware placement remains critical across architectures: when all-to-all exchanges spill outside high-bandwidth NVLink domains into the slower scale-out fabric, workloads suffer sharp bandwidth drops, a placement sensitivity related to the torus geometry penalties observed here. C. Implications for Workload Management These findings indicate that job placement should be treated as a co-design problem rather than purely as a utilization problem. Current allocation policies appear to treat the padding of large jobs (12.5% for our 1,024-node submissions) as unmanaged overhead and, consistent with maximizing occupancy, often place such jobs in elongated remnants. However, for distributed quantum simulation, converting that same

padding into an explicit geometry request improves performance by a factor of two—well above the 20–30% runtime improvement that Pascual et al. found sufficient to compensate for contiguous allocation in fat-tree simulations [21]. This suggests three operational guidelines for workload managers. First, submission interfaces should support a shape intent specification, at minimum a flag indicating preference for compact geometry, enabling schedulers to balance node utilization against predicted execution slowdowns. Second, schedulers can incorporate the surrogate model in Eq. (1) at queue evaluation time to quantify the trade-off between immediate execution on flat regions and waiting for compact allocations. Third, because energy falls together with execution time on compact partitions (flat partitions consume 1.82× more energy on average, Section IV-E), compact placement improves both performance and energy efficiency, so the two objectives need not be traded off in this workload. Implementing these policies requires no application modifications and can be integrated directly into the Quantum–HPC middleware layer discussed in Section II. Regarding the 128 idle nodes incurred under partial use (Section III-E), these nodes belong to the job’s node-exclusive allocation and cannot be released to other users under the current scheduler. The job itself can still use them: Fujitsu MPI supports simultaneous mpiexec launches on disjoint node sets within one job. Work that generates little network traffic, such as single-node circuit generation or result post-processing, is the natural fit; multi-node tasks are possible but would share Tofu links with the simulation and risk slowing it. VI. C ONCLUSION When its node count is set by memory, distributed quantum circuit simulation is a network-bound workload whose execution performance is strongly affected by placement decisions fixed at job submission. Across 764 controlled configurations on Fugaku, we demonstrated that these placement decisions follow a clear hierarchy: partition geometry acts as a consistent factor-of-two performance lever (1.73–2.31× across all twelve algorithms under invariant traffic volume), whereas rank density acts as a secondary, conditional lever shortening runtime by 11% only on compact geometries. The geometry effect can be modeled to first order using a scalar predictive metric, the partition aspect ratio, with runtime scaling as α0.39 (r = 0.72). The resulting best practice, requesting a near-cubic partition with 16 ranks per node, improves time-to-solution by roughly 2× relative to the flat partitions that default placement mostly assigns, which also consume 1.82× more energy than nearcubic ones. More broadly, these findings indicate that partition geometry should be managed as a primary scheduling parameter rather than an unmonitored artifact of utilization packing. Future work should establish formal confidence intervals through repeated trials, validate the placement hierarchy and powerlaw scaling on GPU-based simulators and alternative torus architectures, and integrate the aspect-ratio surrogate into production workload managers for cost-aware scheduling.

The performance penalty of ignoring geometry is severe, the predictive model requires only a single parameter, and the mechanism to request controlled placement can be deployed on unmodified production schedulers today, provided the allocated shape is checked at job start. ACKNOWLEDGMENT First and foremost, we sincerely thank the anonymous reviewers for their in-depth comments. The authors used Anthropic’s Claude as a writing and engineering aid when preparing this manuscript. The tool was used to proofread and polish all sections of this manuscript, to write the scripts that generate Figures 2-5 and Tables III-V from the experimental results. All experimental data, findings and conclusions are from the authors and the authors reviewed, verified, and take full responsibility for all content. This work used computational resources of the supercomputer Fugaku provided by RIKEN Center for Computational Science (Project IDs: ra260020 and ra250027) and was partially supported by Japan Science and Technology Agency (JST) through the Program on Open Innovation Platforms for Industry-academia Co-creation (COINEXT, Grant No.: JPMJPF2221). The MPI version of Qulacs was provided by Fujitsu Research, Fujitsu Ltd. R EFERENCES [1] B. Villalonga et al., “A flexible high-performance simulator for verifying and benchmarking quantum circuits implemented on real hardware,” npj Quantum Information, vol. 5, no. 1, p. 86, 2019. [2] T. Häner and D. S. Steiger, “0.5 petabyte simulation of a 45-qubit quantum circuit,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’17). ACM, 2017, pp. 1–10. [3] A. Tabuchi et al., “mpiQulacs: A scalable distributed quantum computer simulator for ARM-based clusters,” in Proceedings of the 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), 2023, pp. 959–969. [4] T. Jones, A. Brown, I. Bush, and S. C. Benjamin, “QuEST and high performance simulation of quantum computers,” Scientific Reports, vol. 9, no. 1, p. 10736, 2019. [5] Y. Suzuki et al., “Qulacs: a fast and versatile quantum circuit simulator for research purpose,” Quantum, vol. 5, p. 559, Oct. 2021. [6] G. G. Guerreschi, J. Hogaboam, F. Baruffa, and N. P. D. Sawaya, “Intel Quantum Simulator: A cloud-ready high-performance simulator of quantum circuits,” Quantum Science and Technology, vol. 5, no. 3, p. 034007, 2020. [7] H. De Raedt et al., “Massively parallel quantum computer simulator, eleven years later,” Computer Physics Communications, vol. 237, pp. 47–61, 2019. [8] H. Bayraktar et al., “cuQuantum SDK: A high-performance library for accelerating quantum science,” in Proceedings of the 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), 2023, pp. 1050–1061. [9] E. Pednault et al., “Pareto-efficient quantum circuit simulation using tensor contraction deferral,” arXiv:1710.05867v4, 2020. [10] C.-C. Wang, Y.-J. Wang, C.-H. Tu, and S.-H. Hung, “Large-scale quantum circuit simulation on HPC cluster via cache blocking, boosting, and gate fusion optimization,” in Proceedings of the 55th International Conference on Parallel Processing (ICPP ’26). ACM, 2026, pp. 89–99. [11] Y. Teranishi, S. Hiraoka, W. Mizukami, M. Okita, and F. Ino, “Lazy qubit reordering for accelerating parallel state-vector-based quantum circuit simulation,” ACM Transactions on Quantum Computing, vol. 6, no. 4, pp. 27:1–27:33, 2025. [12] X.-C. Wu et al., “Full-state quantum circuit simulation by using data compression,” in Proc. International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2019, pp. 1–24.

[13] Z.-Y. Chen, Q. Zhou, C. Xue, X. Yang, G.-C. Guo, and G.-P. Guo, “64qubit quantum circuit simulation,” Science Bulletin, vol. 63, no. 15, pp. 964–971, 2018. [14] W. M. Brown, A. Ramesh, T. Lubinski, T. Nguyen, and D. E. Bernal Neira, “Multi-GPU quantum circuit simulation and the impact of network performance,” Computer Physics Communications, vol. 324, p. 110126, 2026. [15] A. Bhatele, K. Mohror, S. H. Langer, and K. E. Isaacs, “There goes the neighborhood: Performance degradation due to nearby jobs,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis (SC ’13). ACM, 2013, pp. 1–12. [16] K. Li, M. Malawski, and J. Nabrzyski, “Topology-aware job allocation in 3D torus-based HPC systems with hard job priority constraints,” Procedia Computer Science, vol. 108, pp. 515–524, 2017, International Conference on Computational Science (ICCS 2017). [17] X. Yang, J. Jenkins, M. Mubarak, R. B. Ross, and Z. Lan, “Watch out for the bully! Job interference study on dragonfly network,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’16). IEEE, 2016, pp. 750–760. [18] A. Jokanovic, J. C. Sancho, G. Rodriguez, A. Lucero, C. Minkenberg, and J. Labarta, “Quiet neighborhoods: Key to protect job performance predictability,” in IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2015, pp. 449–459. [19] D. De Sensi, S. Di Girolamo, K. H. McMahon, D. Roweth, and T. Hoefler, “An in-depth analysis of the Slingshot interconnect,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’20). IEEE, 2020, pp. 1–14. [20] Y. Ajima et al., “The Tofu interconnect D,” in Proc. IEEE International Conference on Cluster Computing (CLUSTER), 2018, pp. 646–654. [21] J. A. Pascual, J. Navaridas, and J. Miguel-Alonso, “Effects of topologyaware allocation policies on scheduling performance,” in Job Scheduling Strategies for Parallel Processing (JSSPP 2009), ser. Lecture Notes in Computer Science, vol. 5798. Springer, 2009, pp. 138–156. [22] Z. Zhou et al., “Improving batch scheduling on Blue Gene/Q by relaxing network allocation constraints,” IEEE Transactions on Parallel and Distributed Systems, vol. 27, no. 11, pp. 3269–3282, 2016. [23] Y. Oltchik and O. Schwartz, “Network partitioning and avoidable contention,” in Proceedings of the 32nd ACM Symposium on Parallelism in Algorithms and Architectures (SPAA). ACM, 2020, pp. 563–565. [24] F. Abecassis, V. Karakasis, B. Nabong, and D. Wightman, “Achieving peak system and workload efficiency on NVIDIA GB200 NVL72 with Slurm block scheduling,” NVIDIA Technical Blog, May 2026, https: //developer.nvidia.com/blog/achieving-peak-system-and-workloadefficiency-on-nvidia-gb200-nvl72-with-slurm-block-scheduling/. [25] Y. Mao, S. Shresthamali, and M. Kondo, “Q-Gen: A parameterized quantum circuit generator,” IEEE Transactions on Quantum Engineering, vol. 6, pp. 1–16, 2025. [26] T. Tomesh et al., “SupermarQ: A scalable quantum benchmark suite,” in Proc. IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2022, pp. 587–603. [27] A. Li, S. A. Stein, S. Krishnamoorthy, and J. Ang, “QASMBench: A low-level quantum benchmark suite for NISQ evaluation and simulation,” ACM Transactions on Quantum Computing, vol. 4, no. 2, pp. 1–26, 2023. [28] P. Mantha, F. J. Kiwit, N. Saurabh, S. Jha, and A. Luckow, “PilotQuantum: A middleware for quantum-HPC resource, workload and task management,” in Proceedings of the 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). IEEE, 2025, pp. 1–10. [29] M. Cipollini et al., “Three ways to share a QPU: Scheduling strategies for hybrid quantum-HPC applications,” Future Generation Computer Systems, vol. 185, p. 108699, 2026. [30] P. Viviani et al., “Assessing the elephant in the room in scheduling for current hybrid HPC-QC clusters,” in Proceedings of the 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W). IEEE, 2025, pp. 184–187. [31] T. Badts et al., “Examining QRMI as a unified interface for quantumHPC integration,” arXiv:2607.19591, 2026.

Record · ID 1108714 · SHA-256 39c8ff177bda77ff
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.