ConceptioArchivearXiv CS
arXiv CSopen access

NUMA balancing hampering performance of spiking network simulations

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

arXiv:2607.22275v1 [cs.DC] 24 Jul 2026

NUMA balancing hampering performance of spiking network simulations Melissa Lober

Alp Inangu

Gorka Peraza Coppola

Institute for Advanced Simulation (IAS-6) Jülich Research Centre RWTH Aachen University Jülich/Aachen, Germany [email protected]

Institute for Advanced Simulation (IAS-6) Jülich Research Centre RWTH Aachen University Jülich/Aachen, Germany

Institute for Advanced Simulation (IAS-6) Jülich Research Centre RWTH Aachen University Jülich/Aachen, Germany

Dennis Terhorst

Sebastian Gillessen

Jan Vogelsang

Institute for Advanced Simulation (IAS-6) Jülich Research Centre Jülich, Germany

Institute for Advanced Simulation (IAS-6) Jülich Research Centre Jülich, Germany

Neuromorphic Software Ecosystems (PGI-15) Jülich Research Centre Jülich, Germany

Hans Ekkehard Plesser

Brian Wylie

Benedikt Steinbusch

Department of Data Science, Faculty of Science and Technology Norwegian University of Life Sciences Institute for Advanced Simulation (IAS-6) Jülich Research Centre Ås, Norway / Jülich, Germany

Jülich Supercomputing Centre Jülich Research Centre Jülich, Germany

Jülich Supercomputing Centre Jülich Research Centre Jülich, Germany

Guido Trensch

Susanne Kunkel

Simulation and Data Laboratory Neuroscience Neuromorphic Software Ecosystems (PGI-15) Jülich Supercomputing Centre Jülich Research Centre Jülich Research Centre Jülich, Germany Jülich, Germany

Abstract—Computing centers today mostly operate conventional CPU- and GPU-based systems, where the direct way of decreasing energy consumption is a reduction in the applications’ runtime. Neuromorphic computing promises an alternative architecture with improved energy efficiency for artificial intelligence. In this endeavor, code for the simulation of large-scale spiking networks on conventional supercomputers is the reference. We show that turning off automatic NUMA balancing may reduce energy consumption by 30 %. This dwarfs other attempts of increasing the energy efficiency of a computing center with respect to cost effectiveness. The memory access pattern of spiking network simulation code dynamically interacts with automatic NUMA balancing. This does not affect the correctness of simulation results and thus goes unnoticed in day-to-day neuroscience research. In performance analysis, however, time measurements fluctuate obstructing attempts to optimize simula-

Markus Diesmann Institute for Advanced Simulation (IAS-6) Jülich Research Centre Department of Physics, Faculty 1 Department of Psychiatry,Psychotherapy and Psychosomatics School of Medicine RWTH Aachen University Jülich/Aachen, Germany

tion technology. A new time- and compute-node resolved performance display exposes the fine-grained temporal variability in the course of distributed spiking network simulations. The analysis uncovers that automatic NUMA balancing is of disadvantage and, in particular, affects the jemalloc library for thread-aware memory allocation in a transient manner. The method also allows developers to detect perturbations of the HPC system and target specific improvements to simulation technology. As a consequence of these findings, we have equipped our supercomputers with an option to turn on or off automatic NUMA balancing on a perjob basis on the user level. This gives researchers the opportunity to find the best setting for the application at hand. There are indications in the literature that the effect has been observed before, yet it does not seem common knowledge in scientific computing. It remains to be investigated how widespread the phenomenon is among scientific codes.

I. I NTRODUCTION

false true deliver spikes update neurons collocate spikes

cycle time

There is a discrepancy of orders of magnitude between the energy consumption of current supercomputers and the human brain. While the hardware substrate is different, the two systems are also optimized for different activation patterns: repetitive dense matrix multiplications on the one side, and sparse spatio-temporal activity on the other. Therefore, the field of neuromorphic computing strives to harness the architectural differences for a reduction of the energy consumption of artificial intelligence. Any relevant neuromorphic design needs to outperform conventional computing systems in either speed, energy consumption, or both. This requires the availability of portable benchmarks and highly optimized simulation code exposing the real hardware limits of conventional systems. A model of the cortical microcircuit [1] became a de facto standard benchmark, and a range of neuromorphic systems pass the benchmark today faster than real time [2]. This benchmark constitutes a relevant milestone for the field as any larger model of the mammalian cortex necessarily has a lower average connection probability and should therefore be easier to simulate. Research is now ongoing towards next generation benchmarks. A promising target is a multi-area model (MAM, [3], [4]), demonstrating the interaction of local and global connectivity. The MAM challenges conventional supercomputers because of its communication demands and neuromorphic systems due to its sheer size. Nevertheless, first results based on GPUs as a bridge technology are coming up [5], [6], [7], [8], and new large-scale neuromorphic systems are under construction [9]. The distributed simulation code NEST [10] for spiking neuronal networks serves as a conventional CPU-based reference for new systems. In every simulation cycle, each MPI process independently performs spike delivery to local target neurons, update of all local neurons, and collocation of locally generated spikes; each cycle terminates with global synchronization and spike communication between MPI processes (Fig. 1). Spike delivery and neuron update are fully thread-parallel, while collocation and communication are executed by the master thread alone. The global synchronization requires all MPI processes to wait for the slowest participating process. The real-time factor (RTF), defined as the ratio of wall-clock time to simulated biological time, therefore accumulates the slowest simulation-cycle times across all processes and all simulation cycles [11]. On conventional supercomputers, neither bandwidth nor latency limit simulations of the MAM but the variability in simulation-cycle times, which translates into long synchronization times [11]. However, the cycle time only weakly depends on computational load such that an imbalance of load does not explain it. This situation motivates us to investigate how much of the variability comes from outside the simulation code and is due to the operating system or the hardware. One such operating system feature is automatic NUMA balancing, which was introduced with full scheduler support in Linux 3.13 [12], [13]. On multi-socket servers, memory attached to

synchronize globally communicate spikes

increment time:

Fig. 1. Simulation cycle of distributed spiking network code. Colors indicate simulation phases; brace defines the phases contributing to the simulation-cycle time.

a remote processor socket incurs higher access latency than local memory. To reduce remote memory accesses, the kernel periodically unmaps memory pages and records which NUMA domain faults on each page, then migrates pages toward the domain that accesses them most often. The kernel documentation notes that the overhead of unmapping and fault handling does not always yield a net performance improvement [14]. In the following we first show that NUMA balancing interferes with the distributed and parallel operation of the NEST simulation code in the presence of thread aware memory allocation. We then confirm that the phenomenon also occurs for the system allocator, and finally that it is not particular to a specific number of compute nodes. The presented conceptual and algorithmic work is part of our long-term collaborative project to provide the technology for neural systems simulations [10]. Preliminary results have been presented in abstract form [15], [16]. II. R ESULTS Fig. 2 shows the distribution of simulation-cycle times pooled across all MPI processes and cycles of a MAM simulation with NEST for four combinations of NUMAbalancing state and memory allocator. With automatic NUMA balancing enabled and jemalloc [17] as the memory allocator (Fig. 2A), the distribution exhibits a pronounced long tail extending to cycle times well above the mean, with a maximum observed cycle time of 222.84 ms and a standard deviation of 0.83 ms. Because the RTF is set by the maximum cycle time encountered across all processes at each cycle, this long tail disproportionately inflates the total wall-clock time. With NUMA balancing disabled and jemalloc active (panel

1.5 1.0 99.66

0.5 0.0

C

NUMA-b.=OFF, with jemalloc min 1.11, mean 2.41, max 99.98, std 0.42 RTF = 26.67

density (1/ms)

2.5 2.0 1.5 1.0

99.97

0.5 0.0

0

1

2 3 4 5 simulation-cycle time (ms)

6

99.86

D

NUMA-b.=OFF, with system malloc min 1.33, mean 2.69, max 119.24, std 0.55 RTF = 29.81

99.97 0

1

2 3 4 5 simulation-cycle time (ms)

d3472472-67c9-4da6-ae0c-f2a51fb0e519

2.0

NUMA-b.=ON, with system malloc min 1.32, mean 2.9, max 273.85, std 0.97 RTF = 36.35

2b149771-dc46-4e71-b858-275e90534c31

density (1/ms)

2.5

B d504c3a8-dc7c-4ef5-8100-c150e930f4ad

NUMA-b.=ON, with jemalloc min 1.2, mean 3.12, max 222.84, std 0.83 RTF = 41.69

147fb6a4-ee3b-4475-9b69-6d542428596c

A

6

Fig. 2. Distribution of cycle times pooled across all MPI processes and simulation cycles. Cycle times in the presence of automatic NUMA balancing (panels A, B) and without (C, D) using jemalloc (A, C) or the system allocator (B, D). Data shows results of simulations of a multi-area cortical network model (MAM, [3], [4]) for 10 seconds of biological time using 16 MPI processes with 64 threads each on 8 compute nodes of JURECA-DC. Panels display density only for cycle times up to 6 milliseconds (value in distribution shows % covered), and titles show mean and standard deviation of the distributions, and the maximum cycle time observed in milliseconds. Superimposed values in the top right corners show ratio of wall-clock time and full stretch of simulated biological time (realtime factor, RTF). Vertical text in right margin of plots reading bottom to top uniquely identifies the data set (UUID).

C), the long tail in the cycle time distribution is eliminated: the maximum observed cycle time drops to 99.98 ms and the distribution becomes concentrated around the mean with the standard deviation narrowing to 0.42 ms. A natural hypothesis is that the long-tail cycle times in Fig. 2A reflect periods of elevated network activity, as higher spike counts increase the computational load of spike delivery and collocation. Fig. 3 tests this hypothesis by plotting cycle time against the spike count of the previous cycle for each MPI process. With NUMA balancing enabled (panel A), the longest cycle times show no correlation with spike count. This rules out network dynamics as the cause of the long tail. The effect therefore originates at the machine level, outside the simulation code itself. In the absence of NUMA balancing, a weak dependence of cycle time on spike count emerges (panel C). While spike count varies by a factor of five, cycle time increases by only 50% over this interval. Assuming a linear relationship between cycle time and spike count, the squared Pearson correlation coefficient indicates that spike count accounts for only 17 % of the total variability in cycle time. The residual variability of cycle time, however, is

independent of spike count and of the same order as the mean such that a broad slightly tilted ellipse dominates the graphs. We recognize this structure also in the presence of NUMA balancing (panel A) for the low cycle times. Nevertheless, a complex additional cloud emerges reaching out to cycle times of three times the mean with decreasing probability. The cloud appears as composed of further ellipses like the main one centered at progressively larger means. The stack of ellipses fuses into a vertical band destroying the correlation. Fig. 4 shows cycle times as a function of both simulation cycles and MPI process, providing time- and process-resolved insight into the origin of the long tail. With NUMA balancing enabled and jemalloc active (panel A), four characteristic patterns are visible. First, cycle times are elevated across nearly all processes during the initial thousands of simulation cycles before gradually subsiding. This phenomenon originates from the Linux kernel’s adaptive NUMA balancing heuristic, as sketched below. Second, even after the transient subsides, a subset of processes exhibits persistently elevated cycle times relative to their peers throughout the remainder of the simulation, producing a two-tone horizontally striped pattern

2 1

RTF = 41.69

0

C 5 4 3 2 1 0

25

50

RTF = 26.67 75 100 125 150 175 200 spike count

102 101 100

NUMA-b.=OFF, with system malloc

147fb6a4-ee3b-4475-9b69-6d542428596c

simulation-cycle time (ms)

D

NUMA-b.=OFF, with jemalloc

6

RTF = 36.35

103

occurrences

3

104

0

25

50

RTF = 29.81 75 100 125 150 175 200 spike count

104 103

occurrences

4

d3472472-67c9-4da6-ae0c-f2a51fb0e519

5

NUMA-b.=ON, with system malloc

d504c3a8-dc7c-4ef5-8100-c150e930f4ad

simulation-cycle time (ms)

6

0

B

NUMA-b.=ON, with jemalloc

2b149771-dc46-4e71-b858-275e90534c31

A

102 101 100

Fig. 3. Correlation of cycle time to spike count. Same arrangement of panels and annotation as in Fig. 2. Data pooled over all MPI processes in histogram with 100 bins along each dimension (no smoothing). Counts normalized to the color range of the perceptually uniform sequential map ’viridis’ using the logarithmic transform LogNorm of Matplotlib [18]. White means no counts in observation interval.

in the heatmap. A third pattern is vertical stripes spanning all processes. These stripes reflect the fluctuations in neuronal activity, where elevated activity causes longer cycle times. The stripes extend across all processes because the roundrobin distribution of neurons over processes homogenizes computational load. Hence, even a fluctuation confined to a single brain area affects all processes equally. Finally, there are occasional cycles in individual processes exhibiting extremely long cycle times. These blips carry a vanishing probability mass, and therefore their contribution to total wall-clock time of the simulation is insignificant. With NUMA balancing disabled, the heatmaps show that both the initial transient and the striped process asymmetry are largely absent and the blips are even more rare (panel C). The resulting reduction in RTF amounts to approximately 30 %. There is a faint horizontal two-tone pattern of stripes in the background. This indicates that the two MPI processes running on a node do not find exactly identical conditions. As the MPI processes use all cores but share the operating system, one of them may be affected by system tasks. Even for the second half of the simulation, disabling NUMA balancing reduces the RTF by 28 %. JURECA-DC compute nodes are configured at NPS-4 (Fig. 5), exposing eight NUMA domains per node in total and four per socket. In our simulations, one MPI process is placed

on each socket, with 64 threads explicitly pinned one-to-one to the 64 physical cores of that socket, fully occupying all four NUMA domains. The Linux kernel’s automatic NUMA balancing mechanism continuously monitors memory access patterns and attempts to migrate memory pages toward the NUMA domain of the core most frequently accessing them, with the goal of reducing remote memory access latency. With four NUMA domains per socket, each hosting 16 of the 64 pinned threads, the balancer has three potential migration targets within the socket alone for any given page. Fig. 6 illustrates how the execution structure of NEST interacts with the Linux kernel’s automatic NUMA balancing on JURECA-DC. In NEST’s simulation cycle, spike delivery and neuron update are fully thread-parallel, with threads distributed across all four NUMA domains of the socket. The subsequent collocation and global communication of spikes, however, are single-threaded and executed exclusively by the master thread, which resides in one specific NUMA domain determined by its core assignment. During collocation and communication, the master thread — the single thread executing these serial phases while all other threads are idle — reads spike-buffer data that was written in parallel by all threads across the four NUMA domains of the socket. The automatic NUMA balancing of the Linux kernel interprets this as a signal that

1

2

3

4 5 6 cycle / 10,000

7

8

9

10

RTF = 29.81 (29.75)

0

1

2

3

4 5 6 cycle / 10,000

7

8

9

10

simulation-cycle time (ms)

d3472472-67c9-4da6-ae0c-f2a51fb0e519

NUMA-b.=OFF, with system malloc

6 5 4 3 2 1 6

simulation-cycle time (ms)

RTF = 26.67 (26.66)

0

RTF = 36.35 (35.28)

D

NUMA-b.=OFF, with jemalloc

0 2 4 6 8 10 12 14

NUMA-b.=ON, with system malloc

2b149771-dc46-4e71-b858-275e90534c31

RTF = 41.69 (36.95)

d504c3a8-dc7c-4ef5-8100-c150e930f4ad

0 2 4 6 8 10 12 14

C

MPI process

B

NUMA-b.=ON, with jemalloc

147fb6a4-ee3b-4475-9b69-6d542428596c

MPI process

A

5 4 3 2 1

Fig. 4. Simulation-cycle time resolved by MPI process and simulation cycle. Cycle times are color coded. Same arrangement of panels and annotation as in Fig. 2. Resolution of horizontal axis in simulation cycles (0.1 ms model time). Pairs of consecutive MPI ranks starting with 0 (vertical) are on same node. Cutoff of cycle time is at 6 ms, all larger values indicated by white. RTF in parentheses refers to second half of simulation.

the accessed pages belong to the domain of the master thread and migrates them there. When the deliver and update phases resume in the next cycle, all threads participate equally and access their data in parallel. However, the pages now reside in the master thread’s domain, which is remote for the 48 threads pinned to the other three NUMA domains. The balancer now interprets this as yet another misplacement and corrects back. This migration oscillates every simulation cycle, repeating continuously throughout the simulation. The oscillation never converges. The regular alternation of the simulation code between a fully parallel and a singlethreaded phase is fundamentally incompatible with what NUMA balancing is designed for: a situation where memory ends up in the wrong place once and stays correctly placed after a single migration. Every migration the balancer performs is undone by the next phase transition, so each page fault and Translation Lookaside Buffer (TLB) shootdown incurs overhead without ever producing a lasting improvement. The transient slowness observed in the heatmaps is explained by the adaptive scan period of the balancer; early in the simulation the kernel scans and migrates frequently, but after detecting that migrations are not converging it reduces its scan frequency. This decreases but does not eliminate the overhead. The persistent inter-process asymmetry, however, remains to be fully understood. As a consequence of the analysis above, we introduce a slurm [19] option allowing users to enable or disable auto-

matic NUMA balancing on a per-node basis at the job level, without requiring system administrator intervention (see IV for details). We next investigate whether the phenomenon is particular to the use of the custom memory allocator jemalloc instead of the system allocator. A thread aware allocator is relevant for the simulation code because network construction (not discussed here) needs to be a parallel activity [20]. Using the system allocator instead of jemalloc in the presence of NUMA balancing still leads to an asymmetric cycle-time distribution but its tail is less pronounced and the mean is lower (Fig. 2B). The respective heatmap Fig. 4B uncovers that the initial slowness disappears, reducing the overall RTF, but the horizontal stripes persist. They appear slightly fainter than in the presence of jemalloc, an impression confirmed by a small reduction in second-half RTF. The structure of the correlation graph remains essentially unaltered (Fig. 3B). As previously observed for jemalloc, turning off NUMA balancing also removes the tail of the cycle-time distribution when using the system allocator (Fig. 2D). The correlation graphs (Figs. 3C and D) as well as the heatmaps (Figs. 4C and D) show similar trends for jemalloc and the system allocator. However, while disabling NUMA balancing substantially reduces overall and second-half RTF for both allocators, jemalloc consistently achieves the lower cycle times. Thus, while NUMA balancing perturbs the memory placement achieved by jemalloc more severely than that achieved by the system

allocator, the best configuration uses jemalloc and abstains from NUMA balancing. To assess whether the phenomenon is specific to a particular problem size or a single random instantiation of the network, Fig. 7 shows strong-scaling results for the MAM with NUMA balancing enabled and disabled across four node counts and three independent random seeds. With NUMA balancing enabled, the RTF is consistently elevated relative to the NUMAbalancing-off condition across all node counts and seeds tested, and the variability between seeds is substantially larger. The effect is therefore a systematic property of the interaction between the code’s execution pattern and the NUMA balancing heuristic. Dual-Socket EPYC7742 Node Socket 0

Socket 1

AMD EPYC 7742 · 64 cores

AMD EPYC 7742 · 64 cores

NUMA 0

NUMA 1

NUMA 4

NUMA 5

16 cores (CCD 0–1) 2 × DDR4 ch. ~64 GB local ~120–135 ns

16 cores (CCD 2–3) 2 × DDR4 ch. ~64 GB local ~120–135 ns

16 cores (CCD 0–1) 2 × DDR4 ch. ~64 GB local ~120–135 ns

16 cores (CCD 2–3) 2 × DDR4 ch. ~64 GB local ~120–135 ns

32 GB

32 GB

32 GB

32 GB

32 GB

32 GB

32 GB

32 GB

ch 0

ch1

ch 2

ch 3

ch 0

ch 1

ch 2

ch 3

xGMI

I/O die Infinity Fabric · mem ctrl

I/O die Infinity Fabric · mem ctrl

~200–250 ns

NUMA 2

NUMA 3

NUMA 6

NUMA 7

16 cores (CCD 4–5) 2 × DDR4 ch. ~64 GB local ~120–135 ns

16 cores (CCD 6–7) 2 × DDR4 ch. ~64 GB local ~120–135 ns

16 cores (CCD 4–5) 2 × DDR4 ch. ~64 GB local ~120–135 ns

16 cores (CCD 6–7) 2 × DDR4 ch. ~64 GB local ~120–135 ns

32 GB

32 GB

32 GB

32 GB

32 GB

32 GB

32 GB

32 GB

ch 4 0

ch 5

ch 6

ch7

ch 4

ch 5

ch 6

ch 7

intra-socket cross-quadrant: ~30–40 ns 256 GB total · NPS-4

intra-socket cross-quadrant: ~30–40 ns 256 GB total · NPS-4

One NEST cycle on one MPI rank (one socket, 4 NUMA domains) deliver + update thread-parallel, all 4 domains

NUMA 0

NUMA 1

NUMA 2

NUMA 3

16 threads

16 threads

16 threads

16 worker threads 16 threads master thread here

reads spike data from all 4 domains

collocate + communicate single-threaded, master only

III. D ISCUSSION The research described here was carried out over a period of more than a year, from the formalization of the phenomenon into a ticket (Jülich Supercomputing Center JSC ticket #10100311) to the solution alone nine month were required (April to December 2025). The starting point were unexplainable fluctuations in simulation times in a project on the optimization of the mapping of a brain-scale neuronal network onto a supercomputer [11]. Initially, we gathered confidence that the effect originates from outside our code by control simulations on similar hardware with an independent configuration and software stack.

master thread (NUMA 3)

next cycle: migrated back

Fig. 5. Schematic of a JURECA-DC compute node (Atos BullSequana X2410) comprising two AMD EPYC 7742 processors. Each socket contains 64 cores organised into eight Core Complex Dies (CCDs), grouped into four I/O-die quadrants. The BIOS is configured at NPS-4 (Non-Uniform Memory Access, 4 domains per socket), exposing eight NUMA domains per node. Each domain spans two CCDs (16 cores) and two DDR4-3200 memory channels, each populated with one 32 GB DIMM, giving 64 GB of local memory per domain and 512 GB per node in total. The two sockets are connected via AMD’s inter-socket xGMI link. Memory access latency increases with topological distance: from approximately 120–135 ns within a domain, to approximately 30–40 ns across quadrants within the same socket, to approximately 200–250 ns across the xGMI link.

The five main findings of our study are: • automatic NUMA balancing may interact badly with a memory access pattern like the one of the NEST code • implementation of timers in the simulation code NEST (from release 3.10) reporting the duration of individual cycles per node • a new diagram (heatmap) collocating the cycle times of all nodes as a function of consecutive simulation cycles • a new diagram showing the variability of cycle times as a function of the workload • a software switch for the queueing system slurm exposing automatic NUMA balancing to the user-level There are clear statements in the literature that automatic NUMA balancing can degrade the performance of HPC applications [21], [22], [23]. However, the NUMA nodes per socket (NPS) setting and automatic NUMA balancing are largely treated independently. Research is going on to further improve the algorithms for automatic balancing, see e.g. [24], [25]. At least in the field of computational neuroscience, there is little awareness of the potential for large gains in energy efficiency by optimizing such settings and, vice versa, the potential for catastrophic performance if they are overlooked. The neuronal network model we use in the present study is well established and has produced neuroscientific insight [3]. At the same time it is still at the edge of what neuromorphic computing systems can cope with today in terms of sheer memory requirements but also communication load [5], [6], [7], [8]. In some of the brain areas of the network model, the synchronization of the neuronal activity is higher than in nature. This stresses any computing system more than necessary. Improved brain-scale models may ease this stress on communication.

NUMA balancer sees: pages touched from NUMA 3 migrates worker pages toward NUMA 3

result: oscillating migration every cycle fault + copy overhead repeats thousands of times no stable placement is ever reached

Fig. 6. Schematic of one simulation cycle as executed by a single MPI process occupying one full socket (four NUMA domains, 64 cores). The deliver and update phase is fully thread-parallel, with 16 threads executing in each of the four NUMA domains. The subsequent collocation and communication phase is single-threaded, executed by the master thread residing in one domain (here NUMA 3). During this phase, the master thread reads spike data written by threads across all four domains.

A 100

NUMA-balancing=ON

real time factor

80 60 40

B

NUMA-balancing=OFF synchronize communicate collocate deliver update state propagation

20 0

4 8 16 32 number of compute nodes

4 8 16 32 number of compute nodes

Fig. 7. Strong-scaling benchmark of the MAM. A: Automatic NUMA balancing ON; B: automatic NUMA balancing OFF, in presence of jemalloc for ground state of network dynamics. Threads and two MPI processes per node as in Fig. 2. Color code of stages of the update cycle (legend) corresponds to Fig. 1. State propagation (pink) indicates contributions to total simulation time not accounted for by sum of stages. Error bars indicate standard deviation over three repetitions (collapsing to horizontal bar in B).

Looking back, we have carried out thousands of neuroscientific production runs since the inception of the JURECA-DC supercomputer in 2020. As the simulation results were always correct and arrived comfortably fast, the performance issue went unnoticed for a long time. Only in a technical project aiming at the improvement of the communication architecture of NEST for brain-scale models we became suspicious about the fluctuations in run time. It remains an open question how wide spread the phenomenon is among scientific codes. The communication pattern of NEST characterized by frequent collective communication between nodes with symmetrical parallel action of all threads in between is certainly extreme. Nevertheless, the opportunity to reduce energy costs of calculations by 30 % just by improving the software stack, in this case by a single switch, seems relevant as other measures, like the improvement of computing hardware or cooling systems, may be more costly. The combination of a supercomputer constructed for research at the frontier of what is possible today and a research code that is continuously in flux is difficult to debug. The present work is not a tutorial on how to track down a problem in this situation, but describes a particular phenomenon using a single system. Briefly, in the best case the problem is clearly on the side of the computing system or on the side of the research code; but which side is it? In our experience it is helpful to maintain access to two supercomputers with independent software stacks despite the additional administrative efforts. As a further measure we suggest that research groups inspect a new system with time-resolved instrumentation of their realworld application code before production starts. For a generic simulation code like NEST serving a range of different projects the effort is manageable as the validation has to be done only once. This underlines the necessity to consider relevant research software as scientific infrastructure [26].

Over a period of several months the consortium has systematically investigated the effects of settings on all levels from hardware configuration to application code. This required a robust collaborative workflow and the built-up of a data base to structure the communication between researchers in different units and of different level of experience. A recently developed benchmarking framework [27] supported the investigation and the experience gathered flows back into the refinement of such concepts. Even in the absence of automatic NUMA balancing, the distribution of simulation-cycle times remains on the order of the mean cycle time and only weekly correlates with workload. As the system needs to wait for the slowest node in every communication step, reducing the variability is an effective way of reducing simulation time. Due to this maximum operation even algorithms that slightly increase mean cycle time while reducing variability may be successful. Further work now needs to consider that on modern compute nodes the simulation code is hybrid: message-passing between nodes but massively multi-threaded within nodes. The nodes require a similar max-operation between threads as between nodes. Hence, strategies like sparing a core for the operating system may slow down a compute node in the completion of a cycle but reduce variability. Understanding whether the variability originates from multi-threading and how it can be reduced demands more fine grained tools and modeling as we employ in the present study. From the computational neuroscience perspective, our study shows that the performance of our simulation technology is still far away from the limits imposed by hardware. Thus we expect that there is room for further improvements. From the angle of science management the study shows how a team of domain experts, HPC experts, and supercomputer operators can solve a problem with likely impact beyond the domain. Decisive ingredients for the success was a collaborative benchmarking workflow but also mutual trust and flexibility on all sides. IV. M ETHODS Simulation engine All benchmarks are performed with NEST 3.10 [28], an open-source simulation code for large-scale spiking neuronal networks. NEST uses hybrid parallelization based on MPI for distributed-memory execution and OpenMP for sharedmemory parallelism. During each simulation cycle, spikes are delivered, neuronal states are updated, and newly generated spikes are collected before a global communication step exchanges spikes between MPI processes to maintain causality (Fig. 1). To evaluate performance, we use NEST’s built-in high-resolution timers to measure the runtime of the individual simulation phases (see next paragraph for details). Reported runtimes are averaged across all MPI processes, and performance is expressed as the real-time factor, defined as the wall-clock time normalized by the simulated model time. For the simulation of the MAM, one simulation cycle advances the model by 0.1 ms.

Cycle-resolved timers Starting with NEST 3.10, the simulation code can be configured (cmake option -Dwith-cycle-timers=ON) to record per-cycle timing information and spike counts, a profiling capability introduced as part of the present work (see Figs. 2-4). Since spike communication cannot begin until all MPI processes have completed their local computations, NEST 3.10 also provides the option to separately quantify synchronization and communication time by inserting an explicit MPI_Barrier immediately before the collective MPI_Alltoall() operation. This feature is enabled when compiling the code with the cmake option -Dwith-mpi-sync-timer=ON. More information on the built-in timers of NEST can be found in the user-level documentation1 . Network model All simulations throughout this study use the multi-area model of the macaque visual cortex (MAM2 ; [3], [4]), a largescale spiking network model comprising one square millimeter of each of the 32 visual cortical areas. Each area is represented by a layered microcircuit of excitatory and inhibitory leaky integrate-and-fire neurons [1]. Experimental anatomical data determine area-specific neuron numbers and connectivity, resulting in heterogeneous network structure and activity across cortical areas. Approximately one third of all synapses are long-range projections between areas, with transmission delays sampled from Gaussian distributions. All benchmarks use the model’s asynchronous ground-state regime, characterized by stable low-rate activity with an average firing rate of approximately 2.5 spikes/s. The shortest synaptic delay in the model is 0.1 ms. Benchmarking system We perform benchmarks on the CPU partition (JURECADC Phase 2) of Jülich Research Centre’s JURECA system [29]. The node architecture is shown in Fig. 5. Each standard compute node is equipped with two 64-core AMD EPYC 7742 processors (128 cores total) at (2.25 GHz) and (512 GiB) of memory. Nodes are interconnected via Mellanox HDR100 InfiniBand. Applications are compiled with GCC and use OpenMPI for distributed-memory communication, OpenMP for shared-memory parallelism, and jemalloc 5.3.0 as the memory allocator. We employ two MPI processes per node, each running 64 OpenMP threads. Each MPI process is pinned to the physical cores of a single CPU socket, ensuring that the two processes occupy disjoint sets of cores (OMP_PROC_BIND=close, OMP_PLACES=threads, where the application is scheduled with srun --cpubind=mask_cpu:0xFFFFFFFFFFFFFFFF, 0xFFFFFFFFFFFFFFFF0000000000000000). Simultaneous multithreading (SMT) is disabled. Benchmarks are performed using the automated benchmarking pipeline 1 https://nest-simulator.readthedocs.io/en/stable/index.html 2 https://github.com/INM-6/multi-area-model

CI-beNNch [27] that utilizes continuous integration principles to enable fully automated, reproducible, and user-independent benchmarking workflows. User-level control of automatic NUMA balancing As a consequence of the findings of this study, we have equipped our supercomputers, including JURECA-DC, with a NUMA balancing switch accessible at the user level via the slurm job scheduler. Automatic NUMA balancing is enabled by default and can be disabled for a given job by adding #SBATCH --numa-balancing=0 to the job script. ACKNOWLEDGMENTS This project received funding from NeuroSys as part of the initiative “Clusters4Future” funded by the Federal Ministry of Education and Research BMBF (03ZU1106CB,03ZU2106CB); the German Research Foundation (DFG) - 368482240/GRK2416 (MultiSensesMultiScales) and 545776403/FOR5880 (Mod4Comp); Joint Lab HiRSE, the Helmholtz Platform for Research Software Engineering - an innovation pool project of the Helmholtz Association; the European Union’s Horizon Europe Programme under the Specific Grant Agreement No. 101147319 (EBRAINS 2.0 Project); the Joint Lab ‘Supercomputing and Modeling for the Human Brain’ (SMHB) of the Helmholtz Association; the Volkswagen Foundation - 0071143 (STRUCTICITY); Melissa Lober received funding from the HIDA-NORA Mobility Program for a research stay at NMBU, Norway. We gratefully acknowledge the Gauss Centre for Supercomputing e.V. (GCS) for funding this project by providing computing time on the supercomputers SuperMUC-NG at Leibniz Supercomputing Centre (LRZ) and HPE Apollo (Hawk) at the High Performance Computing Center Stuttgart (HLRS), and computing time granted by the JARA Vergabegremium and provided on the JARA Partition part of the supercomputer JURECA at Forschungszentrum Jülich (computation grant JINB33). We thank the members of the NEST Initiative for providing a forum for discussion and continuous support. C ONFLICT OF I NTEREST S TATEMENT The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. DATA AVAILABILITY STATEMENT Source code, simulation and analysis scripts are openly available at Zenodo [30].

R EFERENCES [1] T. C. Potjans and M. Diesmann, “The cell-type specific cortical microcircuit: Relating structure and activity in a full-scale spiking network model,” Cereb. Cortex, vol. 24, no. 3, pp. 785–806, Mar. 2014. [Online]. Available: https://doi.org/10.1093/cercor/bhs358 [2] J. Senk, A. C. Kurth, S. Furber, T. Gemmeke, B. Golosio, A. Heittmann, J. C. Knight, E. Müller, T. Noll, T. Nowotny, G. Peraza Coppola, L. Peres, O. Rhodes, A. Rowley, J. Schemmel, T. Stadtmann, T. Tetzlaff, G. Tiddia, S. J. van Albada, J. Villamar, and M. Diesmann, “Constructive community race: full-density spiking neural network model drives neuromorphic computing,” Neuromorphic Computing and Engineering, vol. 6, no. 1, p. 012001, feb 2026. [Online]. Available: https://doi.org/10.1088/2634-4386/ae379a [3] M. Schmidt, R. Bakker, K. Shen, G. Bezgin, M. Diesmann, and S. J. van Albada, “A multi-scale layer-resolved spiking network model of resting-state dynamics in macaque visual cortical areas,” PLOS Comput. Biol., vol. 14, no. 10, p. e1006359, 2018. [Online]. Available: https://doi.org/10.1371/journal.pcbi.1006359 [4] M. Schmidt, R. Bakker, C. C. Hilgetag, M. Diesmann, and S. J. van Albada, “Multi-scale account of the network structure of macaque visual cortex,” Brain Struct. Funct., vol. 223, no. 3, pp. 1409–1435, Apr. 2018. [Online]. Available: https://doi.org/10.1007/s00429-017-1554-4 [5] G. Tiddia, B. Golosio, J. Albers, J. Senk, F. Simula, J. Pronold, V. Fanti, E. Pastorelli, P. S. Paolucci, and S. J. van Albada, “Fast Simulation of a Multi-Area Spiking Network Model of Macaque Cortex on an MPI-GPU Cluster,” Front. Neuroinform., vol. 16, p. 883333, Jul. 2022. [Online]. Available: https://doi.org/10.3389/fninf.2022.883333 [6] B. Golosio, G. Tiddia, J. Villamar, L. Pontisso, L. Sergi, F. Simula, P. Babu, E. Pastorelli, A. Morrison, M. Diesmann, A. Lonardo, P. Stanislao Paolucci, and J. Senk, “Scalable construction of spiking neural networks using up to thousands of gpus,” Neuromorphic Comput. Eng., vol. 6, no. 2, p. 024012, May 2026. [Online]. Available: http://dx.doi.org/10.1088/2634-4386/ae65d2 [7] J. C. Knight and T. Nowotny, “Larger GPU-accelerated brain simulations with procedural connectivity,” Nat. Comput. Sci., vol. 1, no. 2, pp. 136– 142, 2021. [8] J. Knight, H. Zhu, and T. Nowotny, “Large-scale, mixed-precision brain simulations on heterogeneous accelerators,” in 2026 34rd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (PDP). IEEE. [9] D. Kudithipudi, C. Schuman, C. M. Vineyard, T. Pandit, C. Merkel, R. Kubendran, J. B. Aimone, G. Orchard, C. Mayr, R. Benosman, J. Hays, C. Young, C. Bartolozzi, A. Majumdar, S. G. Cardwell, M. Payvand, S. Buckley, S. Kulkarni, H. A. Gonzalez, G. Cauwenberghs, C. S. Thakur, A. Subramoney, and S. Furber, “Neuromorphic computing at scale,” Nature, vol. 637, no. 8047, pp. 801–812, Jan. 2025. [Online]. Available: https://www.nature.com/articles/s41586-024-08253-8 [10] M.-O. Gewaltig and M. Diesmann, “NEST (NEural Simulation Tool),” Scholarpedia J., vol. 2, no. 4, p. 1430, 2007. [Online]. Available: https://doi.org/10.4249/scholarpedia.1430 [11] M. Lober, M. Diesmann, and S. Kunkel, “Exploiting network topology in brain-scale simulations of spiking neural networks,” Neuromorphic Computing and Engineering, vol. 6, no. 2, p. 024024, jun 2026. [Online]. Available: https://doi.org/10.1088/2634-4386/ae762e [12] J. Corbet, “NUMA scheduling progress,” Linux Weekly News, Oct. 2013, accessed: 2026-07-23. [Online]. Available: https://lwn.net/ Articles/568870/ [13] M. Gorman, R. van Riel, I. Molnar, and P. Zijlstra, “Merge tag ‘balancenuma-v11’: Automatic NUMA balancing,” Linux kernel git repository, commit 3d59eebc5e13, 2013, merged into Linux 3.13. Accessed: 2026-07-23. [Online]. Available: https://git.kernel.org/pub/ scm/linux/kernel/git/torvalds/linux.git/commit/?id=3d59eebc5e13 [14] The Linux Kernel Development Community, “Documentation for /proc/sys/kernel/,” The Linux Kernel Documentation, 2024, accessed: 2026-07-23. [Online]. Available: https://www.kernel.org/doc/ html/latest/admin-guide/sysctl/kernel.html [15] M. Lober, A. Inangu, G. Peraza Coppola, D. Terhorst, S. Gillessen, J. Vogelsang, H. E. Plesser, S. Kunkel, B. Wylie, B. Steinbusch, G. Trensch, and M. Diesmann, “Numa balancing hampering performance of spiking network simulations,” in NEST Conference 2026, 2026. [Online]. Available: https://events.hifis.net/event/3408/ attachments/5004/12217/NEST-Conference-2026-booklet.pdf

[16] M. Lober, A. Inangu, G. Peraza Coppola, D. Terhorst, S. Gillessen, J. Vogelsang, H. E. Plesser, B. Wylie, B. Steinbusch, G. Trensch, S. Kunkel, and M. Diesmann, “Numa balancing hampering performance of spiking network simulations,” in ICNCE 2026, 2026, p152. [17] J. Evans, “A Scalable Concurrent malloc(3) Implementation for FreeBSD,” in Proceedings of the BSDCan Conference, April 2006. [Online]. Available: https://people.freebsd.org/∼jasone/jemalloc/ bsdcan2006/jemalloc.pdf [18] J. D. Hunter and the Matplotlib development team, “Matplotlib: Visualization with Python,” https://matplotlib.org, 2007, version 2.0+ with default colormap viridis. [19] M. A. Jette, A. B. Yoo, and M. Grondona, “SLURM: Simple linux utility for resource management,” in Proceedings of the 9th International Workshop on Job Scheduling Strategies for Parallel Processing, ser. Lecture Notes in Computer Science, vol. 2862. Springer, 2003, pp. 44–60. [20] T. Ippen, J. M. Eppler, H. E. Plesser, and M. Diesmann, “Constructing neuronal network models in massively parallel environments,” Front. Neuroinform., vol. 11, p. 30, 2017. [Online]. Available: https: //www.frontiersin.org/article/10.3389/fninf.2017.00030 [21] F. Gaud, B. Lepers, J. R. Funston, M. Dashti, A. Fedorova, V. Quéma, R. Lachaize, and M. Roth, “Challenges of memory management on modern NUMA systems,” Communications of the ACM, vol. 58, no. 12, pp. 59–66, 2015. [Online]. Available: https://dl.acm.org/doi/10.1145/2814328 [22] AMD, Inc., “High performance computing tuning guide for AMD EPYC™ 9004 Series Processors,” Advanced Micro Devices, Inc., Technical Report, 2023, aMD EPYC 9004 Series Tuning Guide. [Online]. Available: https://docs.amd.com/api/khub/documents/ NgOfoW49HKdzztTeLBbekA/content [23] M. Gorman and M. Jambor, “Optimizing linux for AMD EPYC™ 9004 Series Processors with SUSE Linux Enterprise Server 15 SP4,” SUSE Software Solutions Germany GmbH, SUSE Best Practices, 2024, accessed: 09 July 2026. [Online]. Available: https://documentation.suse.com/sbp/tuning-performance/pdf/ SBP-AMD-EPYC-4-SLES15SP4 en.pdf [24] F. Gaud, B. Lepers, J. R. Funston, M. Dashti, A. Fedorova, V. Quéma, R. Lachaize, and M. Roth, “Challenges of memory management on modern n-UMA systems,” Communications of the ACM, Practice Section, Aug. 2023, online article. [Online]. Available: https://cacm.acm.org/practice/ challenges-of-memory-management-on-modern-n-uma-systems/ [25] J. Liu and Z. Yu, “Global-state aware automatic NUMA balancing,” in Proceedings of the 15th Asia-Pacific Symposium on Internetware (Internetware 2024). New York, NY, USA: Association for Computing Machinery, 2024. [Online]. Available: https://dl.acm.org/doi/10.1145/ 3671016.3671380 [26] A. Hocquet, F. Wieber, G. Gramelsberger, K. Hinsen, M. Diesmann, F. Pasquini Santos, C. Landström, B. Peters, D. Kasprowicz, A. Borrelli, P. Roth, C. A. L. Lee, A. Olteanu, and S. Böschen, “Software in science is ubiquitous yet overlooked,” Nature Computational Science, vol. 4, no. 7, pp. 465–468, Jul. 2024. [27] J. Vogelsang, M. Lober, C. M. Schöfmann, J. Villamar, D. Terhorst, J. Senk, H. E. Plesser, M. Diesmann, S. Kunkel, and A. C. Kurth, “Continuous benchmarking: Keeping pace with an evolving ecosystem of models and technologies,” 2026. [Online]. Available: https://arxiv.org/abs/2604.15919 [28] C. M. Schöfmann, A. Benelhedi, J. Mitchell, S. Spreizer, P. Nagendra Babu, N. Haug, A. Morrison, D. Terhorst, A. Inangu, R. de Schepper, C. Linssen, J.-E. W. Skaar, S. Kunkel, J. M. Eppler, A. Korcsak-Gorzo, M. Lober, J. Vogelsang, G. Trensch, S. Graber, and H. E. Plesser, “NEST 3.10,” 2026. [Online]. Available: https://doi.org/10.5281/zenodo.18268837 [29] P. Thörnig, “JURECA: Data centric and booster modules implementing the modular supercomputing architecture at Jülich Supercomputing Centre,” Journal of large-scale research facilities JLSRF, vol. 7, Oct. 2021. [Online]. Available: https://doi.org/10.17815/jlsrf-7-182 [30] M. Lober, A. Inangu, G. Peraza Coppola, D. Terhorst, S. Gillessen, J. Vogelsang, H. E. Plesser, B. Wylie, B. Steinbusch, G. Trensch, S. Kunkel, and M. Diesmann, “Data for NUMA balancing hampering performance of spiking network simulations,” 2026. [Online]. Available: https://doi.org/10.5281/zenodo.21533544

Record · ID 405625 · SHA-256 bed63fa8e819ecbd
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.