<Society logo(s) and publication title will appear here.> Received 13 May, 2026; revised ; accepted 22 June, 2026; date of current version 16 July 2026. Digital Object Identifier 10.1109//OJCAS.2026.3714974
The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing †
†
†
†
Stefan Scholze1,3 , Johannes Partzsch1,2 , Sebastian Höppner1 , Florian Kelber1 , † † Andreas Dixius1 , Marco Stolba1 , Sirine Arfa1 , Marc Berthel1 , Georg Ellguth1 , Jim Garside4 , Hector A. Gonzalez1,3 , Stephan Hartmann1 , Thomas Kiel-Hocker1 , Dongwei Hu4 , Matthias Jobst1,2 , Khaleelulla Khan Nazeer1,3 , Tim Langer1,2 , Chen Liu1 , Gengting Liu4 , Matthias Lohrmann1 , Mantas Mikaitis4,5 , Felix Neumärker1 , Amirhossein Rostami1,3 , Stefan Schiefer1,3 , Tilo Schubert1 , Delong Shang4 , Bernhard Vogginger1,3 , Yexin Yan1,3 , Steve Furber1 , and Christian Mayr1,2,3
arXiv:2607.24396v1 [cs.ET] 27 Jul 2026
1
Chair of Highly-Parallel VLSI-Systems and Neuro-Microelectronics, Technische Universität Dresden, Germany 2 Centre for Tactile Internet with Human-in-the-loop (CeTI) 3 Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI Dresden) 4 Advanced Processor Technologies Group, Department of Computer Science, University of Manchester, UK 5 now with School of Computer Science, University of Leeds, Leeds, UK † These authors contributed equally to this work. Corresponding author: Johannes Partzsch (email: [email protected]).
This work is partly funded by the German Federal Ministry of Research, Technology and Space (BMFTR) in DAAD project 57616814 (SECAI), by BMFTR and the Free State of Saxony within ScaDS.AI center of excellence for AI research, by the German Federal Ministry for Economic Affairs and Energy (BMWE) under contract 01MN23004F (ESCADE), and by DFG (German Research Foundation) as part of Germany’s Excellence Strategy – EXC 2050/1 and EXC 2050/2 – Project ID 390696704 – Cluster of Excellence “CeTI” of TU Dresden. This work received funding from EFRE (European Regional Development Fund) and the Free State of Saxony under grant 100373652 (SpiNNcloud). The work was funded by the EC Horizon 2020 Framework Programme under grant agreements 720270 (HBP SGA1), 785907 (HBP SGA2), 945539 (HBP SGA3) and Horizon Europe grant agreements 101147319 (EBRAINS 2.0) and 101120727 (PRIMI). The authors gratefully acknowledge the computing time on the high-performance computer at the NHR Center of TU Dresden. This center is jointly supported by BMFTR and the state governments participating in the NHR (www.nhr-verein.de/unsere-partner).
ABSTRACT In deep learning, efficiency gets more and more important to compensate for the ongoing
growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the brain for achieving unprecedented energy efficiency. However, demonstrations of these gains only recently began to grow in complexity and realworld applicability. With SpiNNaker2, we present a chip that bridges the gap between deep networks and neuromorphic computing and allows for flexible exploration of computing approaches that combine both worlds. It features 152 processing elements equipped with an ARM M4F processor and dedicated accelerators, an extended SpiNNaker routing fabric for scalable event-based communication and a range of external interfaces for system integration, including Gbit Ethernet and an LPDDR4 memory interface. We demonstrate performance and efficiency of the SpiNNaker2 chip for neuromorphic and deep network workloads, as well as novel event-based computing approaches. For deep network workloads, the chip achieves up to 4.5 TOPS in high performance mode and up to 2.7 TOPS/W efficiency in high efficiency mode for INT8 workloads. The chip supports spiking neural networks with >150 000 neurons and > 1.8 billion synaptic events/s when simulated with a 1 ms time step. Its low baseline power of less than 250 mW allows for efficiency even under varying workload conditions, allowing to explore sparse and event-based modes of computation. All this demonstrates the chip’s capabilities as a universal hardware platform for scalable brain-inspired computing and its combinations with mainstream deep network approaches.
I. Introduction
The power draw caused by artificial intelligence (AI) applications is growing drastically due to increased usage demands. To counterbalance this trend, more power-efficient
hardware platforms are needed. Neuromorphic computing has been promising a boost in energy efficiency by taking over working principles of the brain. Those include eventbased computation, locality, and co-location of memory
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ VOLUME ,
1
:
and processing. However, demonstrations of neuromorphic hardware are still lacking behind mainstream AI [1], [2]. On the other hand, recent developments in machine learning for more efficient processing, such as million mixtureof-experts in large language models [3], [4] or state-space models [5], move towards similar modes of operation as used by neuromorphic computing. GPUs are not well suited for executing such types of algorithms due to the required finegrained parallelism and irregular workload patterns. Even novel deep learning architectures are not designed to leverage such dynamically sparse computation [6]–[8]. In parallel, the field of neuromorphic computing is reaching a plateau of maturity. Dedicated neuromorphic hardware systems, implementing spiking neural networks (SNN) have been significantly advanced in terms of size [9], efficiency [10]–[13] and usability [14]. Moreover, platforms for hybrid execution of spiking and artificial neural networks have been developed and successfully deployed [15]–[17]. Among the larger-scale neuromorphic systems, SpiNNaker [18] stands out with its focus on scalable event communication [19] and its general programmability. Unlike other neuromorphic hardware whose architecture follows a synapse-neuron processing scheme, SpiNNaker is completely agnostic to the algorithm or neural network model by realizing them in software on microcontroller cores. In this paper, we present the SpiNNaker2 chip, which follows the concept of the established SpiNNaker system, but significantly extends it in its computational capabilities and circuit support for efficient realization of event-based algorithms. Compared to the SpiNNaker chip, the number of processing elements (cores) is significantly increased from 18 to 152. Each core now includes dedicated accelerators for often-used functions, such as matrix multiplications, convolutions, as well as exponentials and logarithms. The chip can adapt flexibly to local workload fluctuations in spiking and event-based neural networks by dynamic voltage and frequency scaling (DVFS) of processing elements [20]. This adaptability is further supported by a low-leakage design implementation strategy, resulting in a low baseline power. With an extended version of the SpiNNaker routing fabric [19] and power-adaptive chip-to-chip links optimized for event-based communication, the SpiNNaker2 chip can be used to build scalable neuromorphic compute platforms up to 10 million cores [21]. In the following, we present the design and implementation approach for the SpiNNaker2 chip and introduce its architecture, giving details on the individual components and the design implementation. We present a range of spiking and deep neural networks, as well as new event-based algorithms, demonstrating the versatility of the chip. II. Design and Implementation Approach A. Design Philosophy
The SpiNNaker2 chip follows key architectural assumptions of the SpiNNaker system [18], further developing 2
and enhancing them for widening the application focus and improving system efficiency. Its core principles are: • Flexibility by implementation of algorithms predominantly in software on microcontrollers. This makes SpiNNaker2 much more versatile than systems dedicated to a class of algorithms, such as those for spiking or deep neural networks. • Distributed processing on comparatively tiny, independently operating processing elements (PE). This allows for integration of many PEs on a chip and flexible distribution of heterogeneous workloads. • Event-based communication together with matching event-based algorithms offers the potential for significantly reduced data transfer. Motivated by spiking neural networks in SpiNNaker, the concept is extended to events with flexible payloads for supporting a wide range of event-based algorithms. • Workload adaptation for highest energy efficiency, by offering a low baseline power and different performance modes through dynamic voltage and frequency scaling (DVFS). With that, algorithmic savings, e.g. in compute, memory access and communication, can be directly translated into energy savings. • Hardware acceleration of the most common operations in neural network based algorithms. • Robust design by employing established digital design implementation techniques, focusing on reliable operation of all system components under the full range of process, voltage and temperature variations. This is a key prerequisite for a usable large-scale system.
The core focus of the SpiNNaker2 chip design is to provide a platform for hardware realization of different flavors and combinations of neural networks, with the potential to accommodate other classes of algorithms that employ sparse, event-based communication and computation. B. Scalability of the System
The SpiNNaker2 chip is designed to be part of a scalable system that can consist of a few chips to many thousands, allowing like SpiNNaker to implement both mobile and mainframe systems. The SpiNNaker routing fabric [19] is the core of this system architecture. It is fitted to neural connectivity in the brain, where one source neuron is connected to many destinations. Each chip is representing a node in a hexagonal grid. The on-chip event router allows for multicasting packets in each node, allowing for an efficient tree-like realization of brain connectivity. Fig. 1 illustrates the routing fabric and an example multi-cast connection. The hexagonal grid could be arbitrarily extended, limited only by the maximum acceptable latency of the system. Messages in the SpiNNaker2 routing fabric carry small payloads or no payload at all, in line with the communication needs of event-based and spiking neural networks. Those types of algorithms profit from the flexibility and reduced VOLUME ,
<Society logo(s) and publication title will appear here.>
FIGURE 1: Hexagonal SpiNNaker communication fabric. The green lines denote an example route of a source address.
located on the left side of the chip. An Ethernet link is provided for high-speed configuration, located on the top right. General purpose interfaces for access to sensors, robots or debugging devices are located in the top left corner in a periphery module. Within the periphery module a processor (ARM M4F) is located with full access to the chip’s memory and configuration bits, used for various control and background tasks. The periphery unit further contains a single ADPLL clock multiplier used for central frequency multiplication, true random number generation and the adaptive body bias generator as further described in Sec. III-C. For addressing reasons but also to ease chip size scalability, we used a tile based implementation and placement flow, which can be seen in the uniform area and macro organization of the chip. The components, the power-management and the design implementation details are explained in the following sections. B. Components 1) Processing Element (PE)
FIGURE 2: SpiNNaker2 chip layout with components
communication load of small packets and a minimal protocol overhead. III. Chip Architecture A. Overview
The layout view of the SpiNNaker2 chip is shown in Fig. 2. The main computation component is the quad unit (Quad), each containing four processing elements (PEs). Event-based communication is realized by the event router, located in the center of the chip for minimized latency from all PEs. There are communication interfaces to neighboring chips for different ranges. Long-range and short-range communication components are located on each edge of the chip for the hexagonal communication grid. Two DRAM interfaces are VOLUME ,
The PE is the basic processing unit (see Fig. 3). It features an ARM M4F core for general purpose execution and accelerators for specific computations. The PE contains 128 kByte SRAM memory for program code and data, which can also be accessed over the Network-on-Chip (NoC, see section 3). Accelerators and NoC interface can run independently of each other. All parts of the PE are interconnected via an AHB/APB bus system and further controlled through 45 different interrupts evaluated by the ARM core. Accelerators are memory-mapped into a joint address space for easy programming from the ARM core. Each PE has a fine-grained DVFS functionality implemented (see III-C). Three different timers are available for applications and time measurements, supplied with a 1 MHz reference clock. Numerical Accelerator: Each PE implements an iterative exponential/natural logarithm accelerator [22]. Applied are s16.15 and s0.31 fixed-point and single-precision floatingpoint formats for both operand and result independently. For the floating-point format, subnormals and out-of-range notifiers (+/-infinity) are supported. For fixed-point formats, the output is saturated to a maximum and minimum value. The precision is configurable by user-programming between 1 and 16 iterations, taking 7 to 22 clock cycles per operation. The module achieves an accuracy between 1-2 units-in-thelast-place (ulp), as well as monotonicity. Random Number Generators: For processes dependent on entropy, the SpiNNaker2 architecture includes Pseudo Random Number Generators (PRNG) in each PE, supplying 32bit per clock cycle, and one global True Random Number Generator (TRNG) [23]. The PRNGs implement the MARS KISS64 algorithm. The TRNG uses a phase frequency detection signal from multiple ADPLL clock generators as a randomization source. The TRNG output can directly be accessed, or sent to the PRNG modules for post-processing. Entropy sources can also be used to scramble the PRNGs. 3
:
2) Quad unit
FIGURE 3: SpiNNaker2 processing element (PE) components
The chip’s PEs are physically and logically grouped into a Quad. Each Quad is responsible for clock and power management of its 4 PEs and provides local management features (hardware semaphores, interrupt requests). PEs can be operated at one of 4 performance levels, a predefined combination of one of two supply voltages and a clock frequency. Frequency scaling is realized by local frequency division of the centrally-generated ADPLL clock. PEs operate in globally asynchronous, locally synchronous (GALS) mode, either individually or as a locally coupled synchronous group within a Quad. PEs within the same Quad can access the full Quad SRAM memory if in synchronous mode.
3) Network on Chip (NoC)
Rounding Accelerator: Each PE is equipped with acceleration of signed/unsigned integer rounding and saturation of 64-, 32- and 16-bit values [24]. Floating-point input can be rounded to BFloat16 output. The available modes are roundto-nearest with up rounding on a tie and stochastic rounding by applying random numbers from the PRNG. The rounding bit position can be configured to round any number of least-significant bits up to 32. Four threads work in parallel supplied by separate PRNG channels, each generating output in 3–4 cycles. Machine Learning Accelerator (MLA): To support fast and energy efficient execution of DNN inference, each PE provides an accelerator for 8bit/16bit signed/unsigned execution of 2D matrix multiplication and 2D convolution [25], [26]. The accelerator consists of a 16x4 output-stationary multiply-and-accumulate (MAC) array and 4 post-processing modules. Each 2 neighboring 8bit MAC cells can be fused together to accept 16bit input column- or row-wise and accumulated in a 24bit register. The output can be quantized via shift and truncation in a post-processing block and written out as 8bit, 16bit or 32bit. Data can be fetched directly over the NoC with prefetch. Further data reuse is achieved by a shift register for the input during 2D convolution. NoC Interface: All PEs can request or send out NoC packet streams in parallel with a DMA module. NoC packets can be simple read/write requests/responses, control packets, exception packets, protocol packets or SpiNNaker2 router packets (see section 4). The event handler processes received SpiNNaker2 packets with two configurable filters and a default handler. It can filter multicast spike packets and store the 32-bit keys into a FIFO buffer mapped to the SRAM. This allows to process received spikes in a batch, e.g. when a certain fill-level of the FIFO is reached. Non-filtered packets are processed by the default packet handler. See supplement for further details about event handler. 4
The NoC serves as the central communication infrastructure, connecting all on-chip components in a 2D mesh topology. It comprises two distinct, interconnected networks: the Data NoC (DNoC) and the Configuration NoC (CNoC) see Fig. 4. The DNoC handles large data transfers, such as DMA operations and spike packets. With a flit size of 192 bits, it enables the transfer of 128 bit data packets in a single transaction, ensuring high throughput. Operating at 300 MHz, the DNoC achieves a bisection bandwidth of 307.2 Gbit/s. Following the GALS design, the clocks of the PEs and NoC routers are asynchronous, interfaced via asynchronous FIFOs. This synchronization introduces a latency of 5 clock cycles per hop. The CNoC is primarily designed for configuration and production testing. It operates with a 32-bit flit size and shares the same packet format as the DNoC. Running at a 100 MHz reference clock, it functions independently of the ADPLL clock. Each Quad unit includes both a DNoC and CNoC router, interconnected via a network bridge. The PEs are connected to the DNoC, while configuration registers are linked to the CNoC. This interconnection allows full access to all network components across both NoCs.
4) Event Router
The event router is responsible for routing multicast event packets, as well as core-to-core, nearest neighbor packets and global read/write packets, shown in Fig. 5. Packets arrive from other chips via the link interfaces and NoC packets from PEs and are presented to the router through its 6 NoC ports. Packets are decoded and arbitrated by a front-end crossbar to different internal engines in parallel. The SpiNNaker2 event router follows the SpiNNaker router concept [19], extending it via more routing entries and more flexible packet types. It operates at a clock speed of 400 MHz. Several packet types with different functionality have been implemented: • Multicast Packet: This packet is provided for communicating event messages to multiple targets in a multichip system [19]. It contains a 32 bit source ID that VOLUME ,
PE
PE
PE
Event
Link
PE
Event
Interface
Link
DRAM
Link
Event
Event
Host
DRAM Interface
Link
Event
Perihery
<
Interface
<Society logo(s) and publication title will appear here.>
Link
Event
Router
Text
FIGURE 6: SpiNNaker2 communication interfaces
Event Link NoC Router
128bit NoC
32bit NoC
FIGURE 4: SpiNNaker2 NoC connections
theoretical peak throughput of 1843.2 Gbit/s to PEs and 460.8 Gbit/s to external links. The SpiNNaker2 event router provides 16 configurable hardware diagnostic counters, conceptually similar to those available in SpiNNaker. These counters enable fine-grained monitoring and classification of specific packet types based on programmable filtering criteria. Details of the internal router architecture and the counters are provided in the Supplementary Material.
FIGURE 5: SpiNNaker2 event router 5) Interfaces
the router uses to forward the packet to one or multiple targets, dependent on the routing entries. Each of the 16K routing entries contains a 32 bit source filter, indicating for each bit either a match or a wildcard, and a set of output links and PEs to forward the packet to. The packet only has 40 bits and can be extended by up to 128 bits of payload, supporting transfer of scalars or short vectors together with the event. • Core-to-Core Packet: This packet is meant for sending a message from one core (PE) in a multi-chip system to one other core in the system, identified by a 16 bit chip address and an 8 bit core address. Up to 128 bit payload can be added to the packet. • Nearest-Neighbour Packet: This packet is provided for sending data to a neighbouring chip or retrieving data from it. It contains a 32 bit address and an optional payload of at most 128 bit. • Global Read/Write Packet: This packet allows direct access to any address on any chip in a multi-chip system by means of a target chip, target core and target address. Access can either be a write or a read request. Read data is sent back to the source. The multicast packet is especially fitted for event-based communication, supporting a tree-like distribution to multiple targets and low number of bits to be transported, minimizing communication load. The multicast router supports a VOLUME ,
The SpiNNaker2 chip provides interfaces to connect external memory and peripherals, an external controller, a host computer and to communicate with other SpiNNaker2 chips. a: LPDDR4
Two LPDDR4 memory interfaces can connect external memory of up to 4 GByte each. The interface contains a Uniquify LPDDR4 PHY and controller operated at 800 MHz clock frequency providing a raw bitrate of 25.6 Gbit/s. A DMA controller for each interface can offload larger data transfers from applications while preventing NoC congestion. b: Event links
Fast and reliable communication is required to transport events between chips. We separate two scenarios: on-board chip to chip connections with a distance of a few centimeters and inter-board communication with distances up to 1.5 m. Both types are available for each link and can be enabled mutually exclusively. Following the different physical requirements, the links make use of different IO standards for most power efficient communication. Each chip offers six event links to use in the hexagonal routing grid, see Fig. 1. Short-range chip-to-chip links: Short-range communication is implemented with a Chiplet-like interface. The eight input-output pad cells use a voltage of only 0.5 V for powerefficient low-swing transmission. Six data lanes with doubledata-rate (DDR) signaling use a 500 MHz differential clock signal (generated by a local ADPLL) with 0.25 V common 5
:
mode, resulting in a transmission speed of 6 Gbit/s. Source synchronous transmission offers advantages for power and robustness here. Each event link has a separate physical transmitter and receiver interface that work independently and can be kept in power saving mode individually. Data transfer makes use of an error detection with 12 bit cyclicredundancy-check (CRC-12) [27]. In case of data loss, an ID field in the packet is used to resend data and to ensure proper data order. Each packet has a size of 96 bit. It is transmitted within two system clock cycles at 125 MHz as there is a serialization factor of 8. For power saving during link idle, clock and data signals are driven to ground, which allows immediate reactivation at next packet request. Long-range board-to-board links: Long range event communication uses a different communication standard, employing two LVDS IO pads per link, one pad per direction. A 1 GHz transmission submits data with DDR at 2 Gbit/s, with the TX clock being generated by a local ADPLL. The clock is recovered at the receiver with clock-data-recovery (CDR). For ensuring the required signal toggles for CDR, the link uses 8b10b coding [28]. As a consequence, the link needs to stay active to keep clock synchronization on the target chip, which allows power saving only with longer lead time. Each short-range and long-range link has several packet counters to monitor event transmission. Details are described in the Supplementary Material. c: Ethernet
The SpiNNaker2 chip provides a 1 Gbit/s UDP Ethernet interface to connect to an external source. Connection to an external PHY is done via SGMII, using its standard frequency of 625 MHz with DDR. Furthermore, to prevent data loss, we extend the UDP protocol by providing UDT support [29]. As a reliable UDP-based data transfer protocol, UDT adds control packets to counteract packet loss, and encompasses congestion control to avoid a transmission jam. As a simpler alternative, we also provide an optional packet counter field after the UDP/UDT header to catch potential packet loss. The Ethernet-to-NoC interface can interpret incoming UDP payloads with specific magic headers as single or multiple NoC packets with dynamic destinations on- or off-chip. It also provides programmable routing LUTs for specific outgoing packet types, outgoing source coordinates, incoming source ports or general raw data. d: GPIO Interfaces
The SpiNNaker2 chip features 27 GPIOs operating at 1.8V, offering a wide range of connectivity options for external sensors, microcontrollers or other peripherals. Specifically, the chip provides: • 4 UART interfaces • QSPI masters with speeds up to 200 Mbit/s • QSPI slaves with speeds up to 200 Mbit/s • I2 C masters with speeds up to 1 Mbit/s • I2 C slave with speeds up to 1 Mbit/s 6
• 13 PWM channels
A flexible GPIO multiplexer allows customization of the GPIO functionality based on the application.
6) Boot and Management
A boot controller has been integrated into the periphery part of the chip, supporting a two-stage boot process: In the first stage, the chip is responsive to one of the management slave interfaces (JTAG, I2 C, SPI) for loading configuration information, selected by two bootstrap pins. Alternatively, it can boot autonomously via reading configuration data from an external SPI flash memory. All chip resources can be accessed and initialized in this stage. Once a high-speed interface has been brought up, the second stage can use it to load the application code. Boot times depend on which chip components need to be initialized and on the size of the application code, being essentially limited by the speed of the external interface used for loading. For helping with boot, control and background tasks, the periphery contains an ARM M4 management processor with 128 kByte memory running at 100 MHz. It is used during boot for speeding up the bring up of external interfaces that require a complex initialization sequence. It can also be used to broadcast application code to all PEs, saving loading time via the external interface. C. Power Management Concept
GlobalFoundries 22FDX is an FDSOI technology which allows for adaptive body biasing (ABB) to compensate for process, voltage and temperature variations for improved speed and reduced power consumption. The SpiNNaker2 chip utilizes Racyics ABX adaptive body biasing technology with the ABB aware implementation approach from [30] that ensures guaranteed speed for the standard cells over all characterized process-voltagetemperature (PVT) corners and in parallel reduces leakage for most efficient operation at different voltages. There are three different power scenarios in the chip. Components necessary for communication and initial configuration are run in a zero-body-biased (ZBB) power domain. This 0.8 V power domain is available directly after power-up and is always-on during chip operation. Power saving mechanisms for this domain are implemented on logic level, like clock gating and multi-bit flops. Each PE runs in a separate switchable power domain. Depending on computational requirements, one can choose between a performance mode at 0.8 V and a low-power mode at 0.5 V. The PE’s clock frequency is switched in parallel to utilize the higher speed of the standard cells at higher supply voltage. As default, PEs are optimized to run with 150 MHz clock frequency at 0.5 V or with 300 MHz at 0.8 V. In addition, the ABX engine applies a forward-bodybiasing (FBB) scheme for optimized leakage at guaranteed speed targets under these ultra-low-voltage conditions [31]. VOLUME ,
<Society logo(s) and publication title will appear here.>
TABLE 1: Implementation details for subcomponents Component
Count
SRAM
Power Domain
Processing element Quad unit DRAM interface
152 38 2
1 Mbit
FBB 0.5 V/0.8 V ZBB 0.8 V ZBB 0.8 V
Event Links Ethernet interface GPIO interface / management Event router Total
6 1 1 1
896 kbit 1.875 Mbit 1 Mbit 2.875 Mbit 158.625 Mbit
ZBB 0.8 V ZBB 0.8 V ZBB 0.8 V ZBB 0.8 V
Each PE can be configured individually for a specific power / speed mode, either by a global controller or by the PE itself. This allows very fine-grained and quick adaptation for most power efficient job execution. We have confirmed with measurements that transient IR drop due to simultaneous power mode changes of PEs does not affect operation, given sufficient supporting capacitors on the PCB. As a fallback, the hardware semaphores in each Quad unit can be used to limit the number of PEs switching at the same time. As most of the chip area is filled with PEs, also most parts of the chip run in ABX controlled power domains. To reduce the number of externally provided different power supplies, all digital circuits run at either 0.5 V or 0.8 V. D. Physical Design Implementation
The SpiNNaker2 chip is realized with 10 independent implementation macros, assembled in a tile based design. Our goal was to prepare a layout that can be flexibly adjusted with macros that have a common size and interface, taking the size of a Quad as reference. All core components are implemented as multiples of this reference size. Only boundary components have an adjusted size related to IO connectivity. Static timing analysis ensures proper functionality in all implementation corners within each macro. All macro interfaces are asynchronous to remove inter-macro timing relations, thus only interface components need to be timing analyzed. This approach allows for scaling of chip size and compute power at reduced tool runtimes and complexity. Synchronous connections are only used for scan connectivity, having relaxed speed requirements and can make use of clock inversion for relaxed hold timing requirements. Macros of identical type are tested in parallel with parallel pattern shifting and results are evaluated on-chip automatically, reducing test time of all 266M gate equivalents significantly. The layout size of SpiNNaker2 chip is 102 mm2 and includes 158.6 Mbit of memory in total. Table 1 shows distribution of on-chip SRAM inside the different components. E. Test Board and Software Stack
Test boards like the one used in this article (see Fig. 12) contain a STM32H743 microcontroller (STM32) in addition to the SpiNNaker2 chip and DRAM. After turning on the VOLUME ,
board, the STM32 automatically boots the SpiNNaker2 chip via SPI. Afterwards, the SpiNNaker2 chip can be controlled via the UDP Ethernet interface and is ready for application. A software stack for SpiNNaker2 was developed consisting of three main components: 1) Chip software: Bare metal ARM programs to run on the SpiNNaker2 PEs are written in C and compiled using a GCC toolchain with custom linker scripts. C-library functions are available for on-chip communication and control of hardware units such as accelerators. 2) A C++ -based low-level host software provides memory-mapped access via UDP to the chip’s register files, SRAM, and DRAM. Further, there are high-level functions to load ARM programs to PEs and for memory block transfers between host and chip. 3) Python-based high-level software for SNN and DNN applications: py-spinnaker2 [32] allows to define and simulate SNNs consisting of neuron populations and projections with static synapses similar to PyNN [33]. Inference of PyTorch or ONNX DNN models can be run using the OctopuScheduler [34], [35] framework. Only a limited number of layer types and operators is currently supported. The SpiNNaker2 Developer Portal at https://spinnaker2. gitlab.io/ provides an overview of the latest software as well as hardware documentation. IV. Results
We have benchmarked the SpiNNaker2 chip in a range of different scenarios, to present a detailed energy and performance profile of the architecture. The chip has been manufactured in GlobalFoundries 22FDX technology and put on a test board (see bottom right of Fig. 12) that allows power measurement on the different voltage domains of the chip separately. All numbers in this section are chip measurements. In case workloads were too short for power measurement, they were repeated periodically and average power extracted. We start with characterizing baseline power and individual accelerators and continue with spiking and deep neural network workloads, as well as more general eventbased algorithms. The code used for the measurements is available at https://gitlab.com/tud-hpsn/public/s2-chip-paper. A. Baseline Power Characterization
We first present the results of profiling the baseline power in 2 different power modes for 3 different scenarios (see Tab. 2). PEs operate at their nominal clock frequencies: 150 MHz at 0.5 V supply voltage or 300 MHz at 0.8 V (see Sec. III-C). If the PEs are off, power and clock are gated. The PEs are in sleep mode if the ARM M4F cores are in their low-power standby state through the WFI (Wait For Interrupt) or the WFE (Wait For Event) instructions. Multiple blocks, like the machine learning accelerator, are clock gated during sleep mode. To measure PE core power without active accelerators, we run the general-purpose coremark algorithm [36] on all 152 PEs in parallel. Midst coremark execution, all ARM M4F cores and SRAM are fully active. 7
:
Corner
PE [mW]
152 PEs [mW]
total [mW]
PEs off PEs [email protected] PEs [email protected]
0.4 1.4 3.6
59.7 214.7 544.0
75.2 235.4 564.7
PEs [email protected] PEs [email protected]
2.7 8.1
403.9 1227.2
424.6 1247.9
1400
CPU (Floating point) Accelerator (Fixed point)
throughput [CoreMark/s]
65968 138168
Time in ns
1000 600
5.0x 5.7x
4.4x
4.4x
5.1x
400 200 0
exp
log
sigmoid
tanh
2.0 1.5 1.0 0.5
conv1
0.0 2.5
fc1
mm1
2.0 1.5 1.0 0.5 0.0
conv2 0
1
2
3
fc2 4
0
1
2
mm2 3
Performance [TOPS]
ARM CMSIS 0.5V (partitioned) ARM CMSIS 0.8V (partitioned)
Accelerator (Floating point) Accelerator (Floating point) *
1200 800
2.5
Energy Efficiency [TOPS/W]
TABLE 2: SpiNNaker2 chip power for idle operation and coremark processing. The first column describes the execution mode and which power lane is supplied to the PEs. The total includes power of periphery, IO and PLL.
4
0
MLA 0.5V (partitioned) MLA 0.8V (partitioned)
1
2
3
4
MLA 0.5V (full utilization) MLA 0.8V (full utilization)
FIGURE 8: Performance and energy efficiency for 6 DNN layers using ARM CMSIS library versus MLA (weights stored in local PE memory). Partitioned refers to layer execution according to a DNN tiling algorithm (see Sec. IVE), full utilization refers to a scenario if all 152 PEs would be used with the same configuration.
softmax
FIGURE 7: Time measurements for the numerical accelerator, running at 150 MHz. (∗ - optimized implementation)
B. Numerical Accelerator
To assess the speed of the numerical accelerator with exp/log execution in software on the ARM core, we perform time measurements for the bare exponential and logarithm, as well as for the activation functions sigmoid, tanh and softmax, which heavily use the exponential function (see Fig. 7). On the ARM core, computing one exponential using the standard C math library in single-precision floating point takes 1060 ns (at 150 MHz clock frequency). The accelerator uses 32 bit fixed point formats as inputs and outputs, so we show results with 32 bit fixed point, giving a speed-up of 5.3x and 4.8x for exponential and logarithm. For the activation functions, we convert single-precision floating-point inputs to 32 bit fixed point before sending them to the accelerator, running at full accuracy. Afterwards, we convert the result back to floating point for the remaining computations on the ARM core. Therefore, we only observe speed-ups of 3.2x-3.5x. An implementation with optimized loop, where processing time of the accelerator is used for post-processing the previous output, improves the speed-up to between 4.4x (tanh and softmax) and 5.7x (exponential). C. Machine Learning Accelerator (MLA)
We benchmark speed and energy efficiency of the machine learning accelerator (MLA) for typical DNN operations and compare it to execution on the ARM processors with the ARM CMSIS library [37]. We measure execution time 8
and energy efficiency of 6 layers selected from ResNet18 [38] and an Open Pre-Trained Transformer (OPT) model [39] (see supplementary material for layer details). Layers are partitioned to fit into the 128 kB SRAM per PE and optimized for MAC cell utilization, assuming that inputs and weights are stored entirely in the on-chip SRAM. All data was averaged over multiple measurements. Fig. 8 shows measurements of ARM core execution, MLA-accelerated execution of partitioned layer tiles and MLA-accelerated execution for full chip utilization (tiled operation on all 152 PEs) for high-throughput (300 MHz clock) and low-power mode (150 MHz clock). Compared to computation on the ARM core, a partitioned layer execution with MLA has a significantly higher throughput (up to 34.3x / 7.6x / 32.0x for convolutional / fully-connected / matrix-matrix-multiplication layers, 300 MHz clock) and is significantly more energy-efficient (up to 22.6x / 6.8x / 13.5x respectively, 150 MHz clock). For layer fc1, the MLA performance is very close to CMSIS, since only 2 out of 152 PEs are used in the partitioned version. In comparison, layer fc2 shows a more noticeable difference due to the usage of 48 PEs. The theoretical maximum throughput of the MLA is 5.837 TOPS, whereas practically 4.563 TOPS (78 % MAC utilization) could be achieved for matrix-matrix multiplications within a Transformer model [39] for full chip utilization (300 MHz clock), see Fig. 8. For fully-connected layers the MAC utilization is only between 16 % and 19 % as only 1 of 4 rows of the MAC array is used. Instead, for convolutional layers the MAC utilization is between 22 % and 50 %, showing a higher overhead for VOLUME ,
<Society logo(s) and publication title will appear here.>
TABLE 3: Comparison of SpiNNaker2 against different other INT8 inference platforms for DNN Platform
Throughput [TOPS]
Power [W]
Efficiency [TOPS / W]
SpiNNaker2 (150 MHz, 0.5 V)
2.281*
0.825*
2.77
SpiNNaker2 (300 MHz, 0.8 V)
4.563*
2.219*
2.06
Greenwaves GAP9 [8] Coral EdgeTPU [8] Intel Mobileye Eye Q5 [8]
0.151 4 12
0.64 2 5
0.34 2.00 2.40
Jetson Orin Nano [40] IBM NorthPole [41], [42]
67 200 820
25 74 300
2.68 2.70 2.73
ARM Ethos N77 [8] 4.1 0.8 (* denote measured values, see Fig. 8)
5.13
max synaptic events per second
Groq Tensor Streaming [8]
1.6 1.4 1.2 1.0 0.8 0.6 0.4 0.2 0.0
At each time step (e.g., every 1 ms), PEs wake up synchronously by a timer interrupt. Each PE first processes the spikes received during the previous time step, retrieving the corresponding target neurons from the 32-bit spike ID and adding the synaptic weights to the neurons’ synaptic input buffers. Then, the neurons are updated. When a neuron spikes, a spike packet (SpiNNaker2 multicast packet with 32bit key) is sent via the event router to other PEs, where it is received by the event handler. Differing from SpiNNaker1, the 32-bit keys are stored to a spike FIFO in SRAM without interrupting the processor, see supplement for detail. PEs process spike packets and neuron updates sequentially, hence the time needed for processing depends on the input spike rate, the connection density and the number of neurons. Note that PEs may get out of phase if the processing demand is too high, hence exact reproducibility is not guaranteed (See supplement for details).
1e7 lif_conv2d lif_neuron
0
200
400 600 number of neurons
800
1000
FIGURE 9: Maximum synaptic events per second per core (at 300 MHz clock) at a 1 ms tick depending on number of target neurons for sparse connectivity (lif_neuron) and convolutional connectivity (lif_conv2d).
configuration and write-back of results to SRAM compared to matrix-matrix multiplication. Although designed mainly for neuromorphic applications, SpiNNaker2’s energy efficiency for deep neural network inference is comparable to the state of practice achieved by other platforms (see Tab. 3), when operators are allocated entirely in the on-chip SRAM. D. Spiking Neural Networks (SNN)
On SpiNNaker2 SNNs are executed in a similar way to SpiNNaker1 [43]: Each processing element simulates a number of neurons and their incoming synapses. The neuron models, such as leaky integrate-and-fire neurons (LIF), are updated in discrete time steps using the Euler method in software running on the ARM core [44]. In this work, all neuron parameters, state variables and synaptic weights are stored in the PE’s SRAM, which limits the size of the networks that can be implemented. This limitation will disappear once the DRAM is supported in the SNN software. VOLUME ,
a: Core Capacity
The number of neurons per core (PE) is an important metric in neuromorphic chips, as it indicates the scale of the overall system. However, this number varies significantly with parameters such as weight precision, logging, connectivity and simulation time-step. To provide a glimpse of the core capacity in the SpiNNaker2 chip, the maximum number of synaptic events that can be maintained at a 1 ms simulation tick are presented in Fig. 9 as a function of the number of neurons. For sparse connectivity and 1024 neurons per core, up to 12 million synaptic events can be processed per second. Conditions of the characterization are given in the Supplementary Material, which also contains a tentative comparison to SpiNNaker. As mentioned, the tradeoffs affecting the number of neurons are numerous. Studying them is out of the scope of this paper. b: Application Example
As an example, we show the simulation of a deep SNN for DVS gesture recognition [45] using the chip’s power management feature (Sec. III-C). In [44] we have demonstrated a complete end-to-end pipeline for deploying deep SNN models on the SpiNNaker2 chip. This includes the training of the SNN in PyTorch, the conversion to the neuromorphic intermediate representation (NIR) [46], and the execution on the SpiNNaker2 chip using post-training quantization. In this work, we focus on the real-time simulation of the SNN (as if events arrive live from a DVS camera) and apply adaptive power management. The SNN contains 5 layers with in total 7859 neurons which are mapped to 57 cores of the chip (additional 64 cores are used for the input spike population), see Fig. 10 and supplement for details. When simulated for 1000 time steps (corresponding to 1 second of the input data) on SpiNNaker2, we achieve a test accuracy of 92.04% which is less than the 94.0% achieved in PyTorch for the full gestures with on average 6 seconds duration. Longer 9
LIF Conv2D LIF
0
0
50
Pop 1 (lif_conv2d)
Pop 2 (lif_conv2d)
SumPool Conv2D
20
Linear
Pop 4 (lif_neuron)
LIF
Pop 5 (lif_neuron)
4 3 2 1 Pop.
0
50
100 time step
150
auto PL 5
4 3 2 1 Pop.
40 20 0
LIF Linear
core
Flatten
high PL 5
20
LIF SumPool
150
40 0
Pop 3 (lif_conv2d)
100 time step
0
50
0.0
0.5
Output
100 time step
150
1.0 1.5 time done [ms]
time done [ms]
Conv2D
Pop 0 (spike_list)
4 3 2 1 Pop.
40
time done [ms]
Input
SpiNNaker2
core
NIR Model
core
low PL 5
time done [ms]
:
1 0
0
200 spikes received
0
200 spikes received
0
200 spikes received
1 0
1 0
2.0
FIGURE 10: SNN for DVS-Gesture recognition using dynamic power management. Left: NIR model of deep SNN and how it is translated to populations on SpiNNaker2. Center: Time needed (“time done”) for processing the incoming spikes and performing neuron updates per core and simulation step. The rows show the results for different performance levels (PLs): low PL (0.5 V, 150 MHz), high PL (0.8 V, 300 MHz), auto PL (automatic selection of PL according to number of received spikes). Right: scatter plot of “time done” vs. number of received spikes per time step for population 3 (all cores, all time steps, blue dots: low PL, red dots: high PL)
simulations were not possible with the current software stack due limited spike storage in SRAM. We apply and compare 3 different power management strategies for the SNN execution: • Low PL: In the low performance level, all cores are operated at 0.5 V and 150 MHz. • High PL: In the high performance level, all cores are operated at 0.8 V and 300 MHz. • Auto PL: Here we apply dynamic voltage and frequency scaling (DVFS) as introduced in [20] to automatically switch between the high and low performance level in each core and time step. The high PL is chosen if the number of received spikes exceeds a user defined threshold.
The results are presented in Table 4, alongside energy measurements across the different performance levels as shown in Fig. 10. When operated at low PL, we used a simulation tick of 3 ms to make sure that the neural processing finishes for all cores. In Fig. 10 (heat map for low PL) one can see that many cores need close to 2 ms to finish the processing (“time done”). Instead, at high PL, a simulation tick of 1 ms 10
is sufficient to finish the processing on time due to the higher core clock frequency. Finally, we apply the automatic PL selection: For this, we extracted from simulations the dependency of the “time done” and the number of received spikes per core (scatter plots in the right column of the figure) and determined a threshold spike count above which the high PL is used. When applying auto PL, the processing always finishes within 1 ms for real-time processing but only operates at the high PL when needed. By this, 28% less energy is consumed compared to the high PL (Tab. 4).
TABLE 4: SNN DVS gesture prediction on SpiNNaker2. Mode
Low PL
High PL
Auto PL
Simulation tick [s] Energy [J]
0.003 2.010
0.001 1.023
0.001 0.741
Accuracy (%)
92.04
92.04
92.04
Further details about the SNN model, the deployment pipeline and comparison to other works can be found in [44]. VOLUME ,
<Society logo(s) and publication title will appear here.>
E. Deep Neural Networks (DNN)
To support the inference of state-of-the-art DNN architectures like convolutional neural networks (CNNs), transformer-based large language models (LLMs) or vision transformers, the scheduling framework OctopuScheduler has been developed (see [34], [35], [47]). It interprets SpiNNaker2 as an MPSoC platform, coordinating the execution of standard DNN layer types such as convolutional layers, fully connected layers, matrix-matrix-multiplications, and others. The scheduling within a layer and for a complete DNN is performed completely on-chip to minimize the communication latency within and between layers. A single scheduler PE controls the start and termination of up to 151 worker PEs within a DNN layer. Following an automatically derived and optimized tiling scheme, the workers perform required DNN layer computations using the PE-local machine learning accelerator (MLA) on the partitioned input and weight tensors, each producing an output tile. The workers generally operate independently and asynchronously within a layer on equally sized tensor tiles, avoiding intra-layer synchronization and thus minimizing the communication overhead between scheduler and worker. For a complete multi-layer DNN model, each layer execution is triggered synchronously by the scheduler (see [35]). We have benchmarked the same 6 DNN layers from Section IV-E at full size using the OctopuScheduler, which includes loading and storing of data to/from DRAM, data preprocessing like data alignment, and the actual execution of the MLA. Execution time, power and energy for each layer are compared to server and edge GPUs (A100 and Jetson Orin Nano) in Tab 5. As shown in Fig. 11, the largest part of the matrixmultiplication layer execution time results from the DRAM transfers of input, weight or output tensors. For convolutional layers, costly off-chip DRAM transfers of intermediate activations can already be avoided by reordering intermediate activations in the on-chip SRAM between two convolutional layers, showcasing the optimization potential for matrixmultiplication layers. The heavy impact of memory transfers further suggests the use of alternative inference paradigms like pipelined execution or depth-first scheduling, which minimize off-chip communication. The worker-scheduler utilization (average worker execution time compared to complete scheduler execution time) ranges from 80.9% to 99.6% for all layers, showing a small scheduling overhead. A detailed description of the DNN scheduling framework OctopuScheduler can be found in [34]. F. Event-Based Algorithms
In the following, we present recently developed event-based algorithms that go beyond classical spiking or deep neural networks, combining the best of both worlds. [48] introduced an event-based version of the gated recurrent unit (GRU), applying a biologically inspired thresholding mechanism to reduce the communication between neurons. [49] demonstrates VOLUME ,
TABLE 5: Measurement results of DNN layer benchmark on SpiNNaker2 (150 MHz clock, including DRAM transfers and scheduling overhead) compared with Nvidia A100 GPU using torchao and Jetson Orin Nano using torch2trt (all INT8). Layer
Platform
Power (W)
Time (ms)
Energy (mJ)
FC 1 FC 1 FC 1
SpiNNaker2 GPU Jetson
0.577 107.1 3.116
0.111 0.012 0.327
0.064 1.332 1.019
FC 2 FC 2 FC 2
SpiNNaker2 GPU Jetson
0.800 108.7 3.707
0.944 0.024 0.403
0.755 2.648 1.494
MM 1
SpiNNaker2
0.612
3.971
2.429
MM 1 MM 1
GPU Jetson
108.0 3.358
0.015 0.347
1.590 1.165
MM 2 MM 2 MM 2
SpiNNaker2 GPU Jetson
0.697 123.2 8.211
24.501 0.034 0.974
17.073 4.176 7.998
CONV 1
SpiNNaker2
0.723
3.035
2.195
CONV 1 CONV 1
GPU Jetson
133.1 5.834
0.053 0.657
7.041 3.833
CONV 2 CONV 2 CONV 2
SpiNNaker2 GPU Jetson
0.597 116.7 4.269
3.918 0.039 0.602
2.339 4.599 2.570
an implementation of the EGRU network on SpiNNaker2 for language modeling and DVS gesture recognition. This implementation demonstrates that significant gains in energy efficiency can be reached versus conventional hardware for single-batch inference: Energy per inference for a language model reduced by a factor of 18 compared to a GPU implementation (65 mJ on SpiNNaker2 vs. 1.19 J on Nvidia A100), with the drawback of an 8x longer execution time. Due to its flexibility, SpiNNaker2 is ideally suited to explore novel learning algorithms: [50] used SpiNNaker2 to train SNNs on-chip with event-based backpropagation (EventProp), which is hard or even impossible on other neuromorphic systems with dedicated digital cores, but easy to realize with SpiNNaker2 as it can send events with flexible payload, e.g., for sending error signals. For the EventProp example in [50], SpiNNaker2 required only 31% of the energy per training step compared to a RTX 4070 GPU. In a similar fashion, [51] applied the biological plausible Eprop learning rule [52] to train recurrent SNN for keyword spotting on a SpiNNaker2 FPGA prototype. We ported the E-prop code to the SpiNNaker2 chip and measured runtime and power: while SpiNNaker2 is significantly slower than a NVIDIA V100 GPU, it requires 8× less energy when considering device utilization, see supplementary section VI-G for details. Recent neuromorphic algorithms for combinatorial 11
:
prepare 28.5% compute 23.1% memory 35.5% scheduling 12.9%
59.3% 13.2% 16.9% 10.6%
9.0% 9.9% 73.9% 7.2%
6.2% 3.0% 71.6% 19.2%
4.0
4.2% 5.2% 84.3% 6.4%
MM 1
MM 2
25
3.5 3.0 2.5 2.0 1.5 1.0
FSM State Execution Time [ms]
0.8 FSM State Execution Time [ms]
FSM State Execution Time [ms]
1.1% 0.5% 98.0% 0.4%
0.6
0.4
0.2
20
15
10
5
0.5 0.0
CONV 1
CONV 2
(a) conv2d LOAD Input 1 EXECUTE
0.0
FC 1
FC 2
(b) vec-mat-mul LOAD Input 2 Other
0
(c) mat-mat-mul
STORE Output Scheduling Overhead
PREPARE
FIGURE 11: Execution times of layers based on (a) 2D convolutions, (b) vector-matrix-multiplications, and (c) matrixmatrix-multiplications at 150 MHz clock. prepare refers to on-chip input reordering between successive layers for (a) and to MLA input alignment for (b, c), compute to execution of the operation kernel using MLA, memory to storage transfers of inputs and outputs, scheduling to overhead of scheduler execution time and synchronization compared to average worker execution time (see proportional shares per layer at top).
optimization also show promising results on SpiNNaker2 [53]. These examples demonstrate how the flexible compute substrate of SpiNNaker2 allows for implementation of new models and for assessing their gains in energy efficiency. G. Integration with Sensors
A single SpiNNaker2 chip can also operate as a standalone system, where its internal PEs interface with external sensors and actuators. The top of Fig. 12 illustrates a single-chip demonstrator, where the testing board serves as the computing platform for a four-wheel-drive rover, which also features a MEMS-based Robosense RS-LiDAR-M1 mounted on top. The example uses one core for parsing sensor data and sending driving commands via UART to the rover, another core for buffering and preprocessing, and a third core that maps driving commands from gestures acquired through a wireless glove with a GRU-based classifier, deployed with a MicroTVM-based flow [54]. The wireless glove interfaces with a daughter board using the “ESPnow” protocol, which forwards data to SpiNNaker2 via I2 C. V. Conclusion
In this paper, we have presented the neuromorphic manycore chip SpiNNaker2. We introduced its architecture and its main components and features. Finally, we demonstrated 12
FIGURE 12: Top: Robotic demonstrator with SpiNNaker2, bottom: mainframe system “SpiNNcloud” with 48-chip board (left) and full installation (right).
its performance for a variety of computational workloads, including spiking neural networks, deep neural networks, as well as event-based computing approaches. While the focus in this paper was on single-chip results, the SpiNNaker2 system architecture has been designed for a wide range of system sizes. As shown in Fig. 12, it is used both in small-scale demonstrators, as well as cloudscale systems, with the current installation at TU Dresden reaching a size of 5 million cores [55], with 35k chips in eight racks. The software stack for large-scale deployment on this system is currently under development. As shown in Tab. 6, SpiNNaker2 increases capacity by almost an order of magnitude compared to its predecessor SpiNNaker, both in terms of on-chip memory and core count, as well as spiking neuron capacity. Moreover, the hardware accelerators available in the SpiNNaker2 PEs can further increase neuron capacity significantly [56], which has not been incorporated in the comparison. Moreover, the numerical accelerators (exp, log, RNG) and the native floating point support in the processors facilitate implementation of learning algorithms and improve throughput, e.g. speeding up event-based weight updates for STDP learning rules via fast exponential computation. Additionally, more flexible packet types simplify communication of learning signals. We are currently working on NESTML [57] support, which will facilitate studies on a great variety of learning models. VOLUME ,
<Society logo(s) and publication title will appear here.>
TABLE 6: Comparison of SpiNNaker2 with SpiNNaker and Loihi/Loihi2 Loihi
Loihi2
SpiNNaker
SpiNNaker2
Technology node Clock frequency Area On-chip SRAM Power Number of cores
14 nm variable 60 mm2 33 MByte < 1.5 W 128
7 nm 1000 Mhz 60 mm2 31 MByte 1.5-2.5 W 128
130 nm 180 MHz 102 mm2 1.8 MByte 0.36-1 W 18
22 nm 150/300 MHz 102 mm2 19.8 MByte 0.24-2.2 W 152
Core features
Neuro-synaptic core with programmable data pipeline
ARM M0 processor
ARM M4F processor, FP32 support, exp/log, RNG and GMM/convolution accelerators Software-based
Learning support
STDP, three-factor learning
Neuro-synaptic core with programmable data pipeline, graded spikes, convolution support STDP, three-factor learning
Routing
mesh core-to-core routing
mesh core-to-core routing
Spiking neuron capacity
131 072
spike
1 048 576
The hardware accelerators in each PE also enable DNN and event-based algorithms to run efficiently on the chip and thus widen the flexibility and applicability of the system. Still, SpiNNaker2 is not a pure DNN inference accelerator. While efficiency for local DNN processing is comparable to more DNN-centered hardware platforms (see Tab. 3), limited DRAM bandwidth often reduces utilization of the DNN accelerators if DNN layers are too big to be stored in on-chip memory (see Fig. 11). Thus, scalability of conventional DNN workloads on a single chip is limited, except for applications where DNN layers stay small and massive parallelization over the input data is required. Bigger DNN models could be spread onto multiple chips via mapping different layers onto different chips and pipelining their execution. By this, DNN size per chip could be reduced to allow more onchip data re-use and limit DRAM data transfer requirements. Moreover, depth-first execution could be utilized to avoid buffering of activations in DRAM between layers. We have started exploring these options. Still, we expect that more sparse and event-based deep neural network algorithms will result in higher efficiency improvements on SpiNNaker2 (see Sec. IV-F). The Loihi/Loihi2 family of systems targets a similar application range, while differing in its architecture from SpiNNaker/SpiNNaker2 by dedicated hardware blocks for synapses and neurons. Together with implementation in a smaller technology node, this allows a higher number of neurons per area. On the other hand, the range of supported algorithms is restricted on Loihi/Loihi2 to SNN and extensions towards event-based algorithms using graded spikes. The software-based approach of the SpiNNaker/SpiNNaker2 systems offers greater flexibility. Moreover, SpiNNaker2 also offers more routing possibilities, with high-speed data transfer on-chip and transmission of data alongside event packets between chips.
VOLUME ,
spike
Software-based Core-to-core and multicast
16 000
Core-to-core and multicast, NoC for on-chip data transfer 150 000
The variety of workloads demonstrates the increased flexibility of the SpiNNaker2 chip for distributed computing with time-varying computing demands. This opens up possibilities for efficiently realizing new types of machine learning algorithms that reduce overall computing and memory demands by allowing more irregular compute and communication patterns. That approach is orthogonal to GPU and classical deep learning accelerators, which achieve their high efficiency from large regular and predictable memory, compute and communication workloads. Saving computations on such hardware platforms on a fine-grained level often does not pay off in terms of energy or performance. In contrast, the SpiNNaker2 chip better allows capitalization on these kind of savings via its low baseline power and hardware support for workload adaptation. At the same time, the SpiNNaker2 chip is designed for real-time operation and integration with sensors and actuators, making it highly suitable for robotics applications. While there are specific hardware blocks for e.g. deep networks or event handling, the chip is suited for any application that can be distributed onto a large number of tiny compute elements that exchange small chunks of data. We are looking forward to explorations of such novel, more efficient algorithms with the SpiNNaker2 chip. REFERENCES [1] C. Ostrau, C. Klarhorst, M. Thies, and U. Rückert, “Benchmarking neuromorphic hardware and its energy expenditure,” Frontiers in Neuroscience, Section Neuromorphic Engineering, vol. 16, 2022. [2] B. Vogginger et al., “Neuromorphic hardware for sustainable ai data centers,” 2024. [3] P. Belcak and R. Wattenhofer, “Exponentially faster language modelling,” 2023. https://arxiv.org/abs/2311.10770. [4] X. O. He, “Mixture of a million experts,” 2024. https://arxiv.org/abs/ 2407.04153. [5] M. Schöne et al., “Scalable event-by-event processing of neuromorphic sensory signals with deep state-space models,” 2024. https://arxiv.org/ abs/2404.18508.
13
:
[6] R. Prabhakar, S. Jairath, and J. L. Shin, “Sambanova SN10 RDU: A 7nm dataflow architecture to accelerate software 2.0,” in IEEE ISSCC, vol. 65, pp. 350–352, 2022. [7] D. Ignjatović, D. W. Bailey, and L. Bajić, “The wormhole ai training processor,” in 2022 IEEE International Solid-State Circuits Conference (ISSCC), vol. 65, pp. 356–358, 2022. [8] A. Reuther et al., “Lincoln AI Computing Survey (LAICS) Update,” in HPEC 2023, pp. 1–7, 2023. ISSN: 2643-1971. [9] S. Schmitt et al., “Neuromorphic hardware in the loop: Training a deep spiking network on the brainscales wafer-scale system,” in International Joint Conference on Neural Networks (IJCNN), 2017. [10] C. Frenkel and G. Indiveri, “Reckon: A 28nm sub-mm2 task-agnostic spiking recurrent neural network processor enabling on-chip learning over second-long timescales,” in International Solid-State Circuits Conference (ISSCC), vol. 65, pp. 1–3, 2022. [11] C. Frenkel, D. Bol, and G. Indiveri, “Bottom-up and top-down approaches for the design of neuromorphic processing systems: Tradeoffs and synergies between natural and artificial intelligence,” Proceedings of the IEEE, vol. 111, no. 6, pp. 623–652, 2023. [12] S. B. Shrestha, J. Timcheck, P. Frady, L. Campos-Macias, and M. Davies, “Efficient video and audio processing with loihi 2,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 13481–13485, 2024. [13] M. Yao et al., “Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip,” Nature Communications, vol. 15, no. 4464, 2024. [14] M. Davies et al., “Advancing neuromorphic computing with loihi: A survey of results and outlook,” Proceedings of the IEEE, vol. 109, no. 5, pp. 911–934, 2021. [15] G. Tang et al., “Seneca: building a fully digital neuromorphic processor, design trade-offs and challenges,” Frontiers in Neuroscience, vol. 17, 2023. [16] J. Pei et al., “Towards artificial general intelligence with hybrid tianjic chip architecture,” Nature, vol. 572, pp. 106–111, 2019. [17] S. Ma et al., “Neuromorphic computing chip with spatiotemporal elasticity for multi-intelligent-tasking robots,” Science Robotics, vol. 7, no. 67, 2022. [18] S. Furber, F. Galluppi, S. Temple, and L. Plana, “The spinnaker project,” Proceedings of the IEEE, vol. 102, no. 5, pp. 652–665, 2014. [19] S. Furber, D. Lester, L. Plana, J. Garside, E. Painkras, S. Temple, and A. Brown, “Overview of the spinnaker system architecture,” IEEE Transactions on Computers, vol. 62, no. 12, pp. 2454–2467, 2013. [20] S. Höppner et al., “Dynamic power management for neuromorphic many-core systems,” IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 66, no. 8, pp. 2973–2986, 2019. [21] C. Mayr, S. Hoeppner, and S. Furber, “SpiNNaker 2: A 10 million core processor system for brain simulation and machine learning,” in Communicating Process Architectures 2017 & 2018, p. 277–280, IOS Press, 2019. [22] M. Mikaitis et al., “Approximate fixed-point elementary function accelerator for the spinnaker-2 neuromorphic chip,” in 2018 IEEE 25th Symposium on Computer Arithmetic (ARITH), pp. 37–44, 2018. [23] F. Neumarker, S. Höppner, A. Dixius, and C. Mayr, “True random number generation from bang-bang adpll jitter,” in 2016 IEEE Nordic Circuits and Systems Conference (NORCAS), pp. 1–5, 2016. [24] M. Mikaitis, “Stochastic rounding: Algorithms and hardware accelerator,” in IJCNN 2021, pp. 1–6, 2021. [25] S. M. A. Zeinolabedin et al., “A 16-channel fully configurable neural soc with 1.52 µw/ch signal acquisition, 2.79 µw/ch real-time spike classifier, and 1.79 tops/w deep neural network accelerator in 22 nm fdsoi,” IEEE TBioCAS, vol. 16, no. 1, pp. 94–107, 2022. [26] F. Kelber et al., “Mapping Deep Neural Networks on SpiNNaker2,” in NICE 2020, pp. 1–3, 2020. [27] T. Ramabadran and S. Gaitonde, “A tutorial on crc computations,” IEEE Micro, vol. 8, no. 4, pp. 62–75, 1988. [28] Advanced Micro Devices, Inc., “Aurora 8B/10B Protocol Specification,” 2014. [29] Y. Gu and R. L. Grossman, “UDT: UDP-based data transfer for high-speed wide area networks,” Computer Networks, vol. 51, no. 7, pp. 1777–1799, 2007. [30] S. Höppner et al., “Adaptive body bias aware implementation for ultra-low-voltage designs in 22fdx technology,” IEEE Transactions on Circuits and Systems II, vol. 67, no. 10, pp. 2159–2163, 2020.
14
[31] S. Höppner et al., “How to achieve world-leading energy efficiency using 22fdx with adaptive body biasing on an arm cortex-m4 iot soc,” in ESSDERC 2019, pp. 66–69, 2019. [32] B. Vogginger et al., “py-spinnaker2,” 2024. https://zenodo.org/doi/10. 5281/zenodo.10202109. [33] A. P. Davison, D. Brüderle, J. M. Eppler, J. Kremkow, E. Muller, D. Pecevski, L. Perrinet, and P. Yger, “Pynn: a common interface for neuronal network simulators,” Frontiers in neuroinformatics, vol. 2, p. 388, 2009. [34] T. Langer, M. Jobst, C. Liu, F. Kelber, B. Vogginger, and C. Mayr, “OctopuScheduler: On-Chip DNN Scheduling on the SpiNNaker2 Neuromorphic MPSoC,” in 2025 Neuro Inspired Computational Elements (NICE), pp. 1–10, Mar. 2025. [35] M. Jobst, T. Langer, C. Liu, M. Alici, H. A. Gonzalez, and C. Mayr, “An end-to-end dnn inference framework for the spinnaker2 neuromorphic mpsoc,” in ICONS 2025, 2025. [36] EEMBC, “Coremark benchmark,” 2025. Accessed: 2025-01-15. [37] L. Lai, N. Suda, and V. Chandra, “Cmsis-nn: Efficient neural network kernels for arm cortex-m cpus,” arXiv:1801.06601, 2018. [38] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” Dec. 2015. arXiv:1512.03385 [cs]. [39] S. Zhang et al., “OPT: Open Pre-trained Transformer Language Models,” 2022. arXiv:2205.01068 [cs]. [40] NVIDIA Corporation, “NVIDIA Jetson Orin Nano Super Developer Kit. Datasheet,” 2024. [41] A. S. Cassidy et al., “11.4 IBM NorthPole: An Architecture for Neural Network Inference with a 12nm Chip,” in 2024 IEEE International Solid-State Circuits Conference (ISSCC), vol. 67, pp. 214–215, 2024. [42] D. S. Modha et al., “Neural inference at the frontier of energy, space, and time,” Science, vol. 382, pp. 329–335, Oct. 2023. [43] O. Rhodes et al., “sPyNNaker: A software package for running PyNN simulations on SpiNNaker,” Frontiers in Neuroscience, vol. 12, 2018. [44] S. Arfa, B. Vogginger, C. Liu, J. Partzsch, M. Schöne, and C. Mayr, “Efficient Deployment of Spiking Neural Networks on SpiNNaker2 for DVS Gesture Recognition Using Neuromorphic Intermediate Representation,” in NICE 2025, pp. 1–8, 2025. [45] A. Amir et al., “A low power, fully event-based gesture recognition system,” in IEEE conference on computer vision and pattern recognition, pp. 7243–7252, 2017. [46] J. E. Pedersen et al., “Neuromorphic intermediate representation: A unified instruction set for interoperable brain-inspired computing,” Nature Communications, vol. 15, p. 8122, Sept. 2024. [47] H. A. Gonzalez et al., “Spinnaker2: A large-scale neuromorphic system for event-based and asynchronous machine learning,” in MLNPCP Workshop at NeurIPS 2023, 2023. [48] A. Subramoney, K. K. Nazeer, M. Schöne, C. Mayr, and D. Kappel, “Efficient recurrent architectures through activity sparsity and sparse back-propagation through time,” in ICLR, 2023. [49] K. K. Nazeer et al., “Language modeling on a SpiNNaker2 neuromorphic chip,” in IEEE AICAS, pp. 492–496, 2024. [50] G. Béna, T. Wunderlich, M. Akl, B. Vogginger, C. Mayr, and H. A. Gonzalez, “Event-based backpropagation on the neuromorphic platform spinnaker2,” in 2025 Neuro Inspired Computational Elements (NICE), pp. 1–10, IEEE, 2025. [51] A. Rostami, B. Vogginger, Y. Yan, and C. G. Mayr, “E-prop on SpiNNaker 2: Exploring online learning in spiking RNNs on neuromorphic hardware,” Frontiers in Neuroscience, vol. 16, 2022. [52] G. Bellec et al., “A solution to the learning dilemma for recurrent networks of spiking neurons,” Nature Communications, vol. 11, no. 1, p. 3625, 2020. [53] Z. Chen et al., “ON-OFF neuromorphic ISING machines using fowlernordheim annealers,” Nature Communications, vol. 16, 2025. [54] C. Liu, M. Jobst, L. Guo, X. Shi, J. Partzsch, and C. Mayr, “Deploying Machine Learning Models to Ahead-of-Time Runtime on Edge Using MicroTVM,” in CODAI Workshop, p. 37–40, 2024. [55] D. Kudithipudi, C. Schuman, C. M. Vineyard, T. Pandit, C. Merkel, R. Kubendran, J. B. Aimone, G. Orchard, C. Mayr, R. Benosman, et al., “Neuromorphic computing at scale,” Nature, vol. 637, no. 8047, pp. 801–812, 2025. [56] J. Huang, F. Kelber, B. Vogginger, C. Liu, F. Kreutz, P. Gerhards, D. Scholz, K. Knobloch, and C. G. Mayr, “Efficient snn multi-cores mac array acceleration on spinnaker 2,” Frontiers in Neuroscience, vol. 17, 2023.
VOLUME ,
<Society logo(s) and publication title will appear here.>
[57] C. Linssen, P. N. Babu, J. M. Eppler, L. Koll, B. Rumpe, and A. Morrison, “Nestml: a generic modeling language and code generation tool for the simulation of spiking neural networks with advanced plasticity rules,” Frontiers in Neuroinformatics, vol. Volume 19 - 2025, 2025. [58] O. Rhodes, L. Peres, A. G. Rowley, A. Gait, L. A. Plana, C. Brenninkmeijer, and S. B. Furber, “Real-time cortical simulation on neuromorphic hardware,” Philosophical Transactions of the Royal Society A, vol. 378, no. 2164, p. 20190160, 2020. [59] J. C. Knight and S. B. Furber, “Synapse-centric mapping of cortical models to the spinnaker neuromorphic architecture,” Frontiers in Neuroscience, vol. Volume 10 - 2016, 2016. [60] L. Peres and O. Rhodes, “Parallelization of neural processing on neuromorphic hardware,” Frontiers in Neuroscience, vol. Volume 16 2022, 2022. [61] A. Rostami, B. Vogginger, Y. Yan, and C. G. Mayr, “E-prop on spinnaker 2: Exploring online learning in spiking rnns on neuromorphic hardware,” Frontiers in Neuroscience, vol. 16, 2022. [62] P. Warden, “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition,” arXiv:1804.03209 [cs], Apr. 2018.
VOLUME ,
15
:
VI. Supplementary Material for the paper ”The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing” A. Internal Structure of Event Router a: Router Architecture
The architecture of the SpiNNaker2 event router is shown in Fig. 13. The router occupies the silicon area of two Quad Processing Elements (QPEs) and is connected to six bidirectional Network-on-Chip (NoC) ports. This configuration enables simultaneous packet reception and transmission across all six ports, providing high aggregate bandwidth. Incoming packets from the six input ports are first processed by an input crossbar interconnect, which performs arbitration and distributes packets to three routing engines operating in parallel. This design enables concurrent processing of heterogeneous packet types, increasing overall throughput and reducing input blocking. Configuration packets are routed directly to the register file module. Each routing engine can independently forward packets to different output ports. Consequently, multiple packets can be routed and issued in parallel within the same cycle, provided that output port conflicts do not occur. This architecture minimizes structural hazards and sustains high throughput under balanced traffic. b: Multicast Routing Engine
FIGURE 13: Internal structure of the SpiNNaker2 event router
Multicast packets are handled by a dedicated multicast routing engine implementing a four-stage pipeline: 1) TCAM Lookup Stage: Performs associative matching of packet address field against routing entries. 2) Priority Encoding Stage: Resolves multiple matches and selects the highest-priority routing entry. 3) Link-Destination Resolution Stage: Determines the set of 7 off-chip links. 4) Core-Destination Lookup Stage: Determines internal 152 core destinations. To improve energy efficiency, the fourth stage can be dynamically disabled when no internal core destinations are present, reducing unnecessary lookup activity and lowering dynamic power consumption. c: Output Scheduling and Contention Handling
To mitigate output contention and avoid pipeline stalls, an out-of-order issue buffer is implemented at the MC routing engine output. If a packet cannot be transmitted due to a blocked output port, it is temporarily stored in the issue buffer rather than stalling the entire pipeline. The scheduler continuously monitors output availability and selects the earliest packet that becomes eligible for transmission. This mechanism provides reduced output blocking, improved output port utilization, higher sustained throughput under contention, resulting in lower average packet latency compared to strictly in-order issue 16
B. Diagnostic Counters for Event Routing
As mentioned in the main text, the SpiNNaker2 event router provides 16 configurable hardware diagnostic counters. Each 32-bit diagnostic counter is associated with a configurable filter control register, allowing selective packet counting according to the following header attributes: 1) Destination Field (8 bits), distinguishing between the seven off-chip links, the monitor core, other cores and dropped packets. 2) Source Field (2 bits), differentiating between off-chip packets and on-chip packets originating from processing elements (PEs). 3) Payload Size Field (4 bits), encoding different packet lengths: packets without payload, packets with 32-bit payload, 64-bit payload or 128-bit payload. 4) Routing Algorithm Field (2 bits), indicating whether default routing was applied or a non-default routing scheme was used. 5) Packet Type Field (3 bits), classifying packets as Nearest Neighbour (NN), Core-to-Core (C2C) or Multicast (MC) packets. In addition to packet classification counters, the SpiNNaker2 router integrates several dedicated hardware profiling counters to support performance analysis and debugging: VOLUME ,
<Society logo(s) and publication title will appear here.>
1) Router Cycle Counter: Measures the number of clock cycles consumed during a specified router operation. 2) Busy Cycle Counter: Counts router wait cycles, i.e., cycles during which backpressure occurs at any of the six router input ports. 3) Zero-Wait Packet Counter: Records the number of packets that traverse one of the MC, C2C, or NN routing engine outputs without incurring any delay. 4) Iterative Drop Counter: Counts packets that are dropped more than once during routing. 5) Error Packet Counter: Records packets associated with routing errors, including time-phase violations and unroutable packet errors. 6) Total Reinserted Packet Counter: Counts the number of packets that are reinserted into the routing fabric by hardware mechanisms. 7) Dropped packet Water-Level Counter: Tracks the number of SpiNNaker packets currently stored in the dropped-packet buffer. The associated water-level register records the maximum occupancy reached in this buffer. C. Diagnostic Counters for Event Links
Each short-range and long-range link has four diagnostic counters for monitoring transmission. The four counters are incremented at the following conditions: 1) Error in CRC of received packet 2) Correct sequence ID in received packet 3) Link partner requested re-send 4) Wrong header in received packet Each counter is 8 bit wide. D. Event Handler and Spike FIFO
As mentioned in Section III.B.1 and Fig. 3, each PE has an event handler to process incoming SpiNNaker packets without interrupting the processor. The event handler has two filters and one default handler. Each filter, also called “spDMA” allows to automatically transfer the content of received SpiNNaker2 packets to an SRAM-mapped FIFO. Each filter is configured with 1 bit per packet type (Multicast, core-to-core or nearest-neighbour) and 1 bit per payload size. There is a match if the type and payload bit corresponding to the received packet are enabled. In addition, one can configure what to store: the control field, the key-field, and payload. The selected data is stored into a hardware FIFO mapped to SRAM. The FIFO size can be set by specifying the start and end SRAM addresses. User code can access and modify the current read and write pointers, read the current fill-level and bits indicating whether the FIFO is full or empty. If the FIFO is full, new packets are either dropped and counted, overwrite existing data, or stall the NoC interface (depending on configuration). While the FIFO is not full, there is no back-pressure on the NoC. Interrupts to the ARM processor can be generated for various cases such VOLUME ,
as the FIFO being non-empty, having reached a certain filllevel or being full. The default handler works in the same way but processes all packets not previously matched by filter 1 or 2. E. Spiking Neural Networks 1) Synchronization
The synchronization of PEs for SNNs works as follows: 1) For each PE, the ARM program is loaded and started to call setup functions according to configuration data stored in SRAM. Then, the PE goes into WFI (wait for interrupt) state. 2) A chip-global interrupt wakes up all PEs simultaneously. Then, the PEs start the timer to trigger a timer interrupt every N reference clock cycles, for example N = 1000 for a 1 ms time step considering the 1 MHz reference clock, cf. Section III.B.1. Again, the PEs go into WFI state. 3) The timer interrupt wakes up all PEs: The neuron cores first read the current spike FIFO’s read and write pointer position to know which spikes were received in the previous time step. Then, the program starts to process these spikes, performs synapse and neuron updates, sends out spikes, and records spikes and voltages (if enabled). Afterwards, the PEs go into WFI state. If in 3) the next timer interrupt is issued before the processing of the previous time step finished, first, the processing of the previous step will run until completion. The processing of the next time step will start delayed so that the spike FIFO may contain spikes received in the new time step, which however will be treated as being from the previous step. Hence, the simulation may get out of phase. Yet, PEs can catch up in subsequent time steps to get in synch again. Such incidents are logged during simulation and reported to the user afterwards. Multi-chip SNN simulation will use the so-called SCAMP program running on each chip to trigger the chip-global interrupt. The SCAMP programs on different chips and boards use a software mechanism to measure and compensate for their relative drift [58].
2) SNN characterization
Details for the study in Fig. 9: We used the high-level library for SNN py-spinnaker2 [32] for the characterization. Two different neuron models are used: lif_neuron – leaky integrate and fire (LIF) neuron with sparse incoming connections (individual synapses stored as list-of-lists similar to sPyNNaker [43]), and lif_conv2d – LIF neuron with 2D convolutional connectivity using shared weights. A spike_list population is used to provide input spikes. The experiment is set up such that the number of synaptic events equal for both neuron models. The 2D convolution has 17
:
TABLE 7: Partitioning of the Deep SNN for DVS-Gesture recognition to SpiNNaker2 PEs (cores) Population
Neurons
Neurons per Core
Nr. of Cores
0 (spike_list) 1 (lif_conv2d) 2 (lif_conv2d) 3 (lif_conv2d) 4 (lif_neuron)
4096 3600 3600 392 256
32 225 225 49 16
64 16 16 8 16
5 (lif_neuron)
11
11
1
1 input layer, 4 output channels and a 3x3 kernel. The size of each output channel is varied between 3x3 (36 neurons) up to 16x16 (1024 neurons). Each spike arriving in one core leads to 4x3x3=36 synaptic events. Then, the number of input spikes per time steps is sweeped and the ”time done” is measured. The maximum number of timesteps is then interpolated. The sweep for the sparse connectivity uses the same neuron numbers as for the 2D convolution, as well as the same number of synaptic events per incoming spikes (36).
3) SNN application
For the Deep SNN in section IV.D we used the following partitioning and placement strategy of populations to SpiNNaker2 PEs: In py-spinnaker2, we manually defined the maximum number of neurons per core (PE) for each population as given in Tab. 7. The software automatically splits larger populations accordingly into multiple population slices and maps them to separate cores on the SpiNNaker2 chip. A simple, linear placement is used: Populations and their slices are placed in provided order to the PEs sorted by Quad coordinate and PE id inside the quad. Here, the number of neurons per core per population was manually adjusted considering memory constraints and to balance computational load.
• Clock frequency: 300 MHz on SpiNNaker2 vs. 200 MHz for SpiNNaker. • Neuron model: A simplified leaky integrate-and-fire neuron without synaptic time constants (SpiNNaker2) vs. slightly more complex leaky integrate-and-fire neuron with exponential synapses (SpiNNaker) • Synapses: 16-bit synapses stored in SRAM (SpiNNaker2) vs. 32-bit synapses stored in SDRAM (SpiNNaker) • Neuron updates: float32 (SpiNNaker2) vs. 32-bit fixed-point (SpiNNaker)
This makes a fair comparison impossible. Instead, we look forward to a future direct comparison between the two systems once the sPyNNaker [43] has been migrated to SpiNNaker2. We also highlight additional work for SpiNNaker splitting synapse and neuron processing to different cores, thereby achieving a higher throughput of synaptic events, see [59], [58], and [60]. F. DNN Layer Parameters
Tables 8 and 9 contain the DNN layer parameters taken from prototype layers of ResNet-18 [38] and an Open PreTrained Transformer (OPT) model [39]. The prototype layers were used for the benchmark of MLA against ARM CMSIS execution (see section IV-C and Fig. 8), focusing only on the layer computation under the assumption that all layer weights and inputs are present locally in SRAM. Additionally, they were used for the full DNN layer benchmark (incl. storage transfers) within the OctopuScheduler framework (see section IV-E, Fig. 11, and Tab. 5). The partitioned layer parameters refer to a distributed layer execution according to the tiling algorithm of OctopuScheduler [34]. TABLE 8: Parameters for complete and partitioned 2Dconvolutional (CONV) layers implemented with OctopuScheduler (a) complete CONV layer parameters
4) Comparison to SpiNNaker
While it would be desirable to provide quantitative numbers on the advancements of SpiNNaker2 over SpiNNaker, a fair comparison of the neuron and synapse processing capacity is currently not possible due to different software stacks. A similar capacity study like the one shown in Fig. 9 has been performed in [59] for SpiNNaker. They report maximum 6.144 million synaptic events per second and core for 256 leaky integrate-and-fire neurons using 100 % connection density. Instead, on SpiNNaker2 we achieve 15.14 million synaptic events at 14 % connection density with 256 neurons per core. Here, SpiNNaker2 is roughly 2.5× better than SpiNNaker. While both approaches use the same format (list of lists) for representing sparse connection matrices, there are the following major differences between implementations: 18
Layer Name
C
Input H W
H
Weights W Strides
C
Output H W
CONV 1 CONV 2
3 64
240 60
7 3
7 3
64 64
120 60
320 80
(2, 2) (1, 1)
160 80
(b) partitioned CONV layer parameters Layer Name
C
Input H W
H
Weights W Strides
C
Output H W
Worker PEs
CONV 1 CONV 2
3 64
13 14
7 3
7 3
32 4
4 12
120 80
166 82
(2, 2) (1, 1)
80 80
The DNN layers were benchmarked on an Nvidia A100SXM4-40GB GPU with CUDA 12.8 using torchao (version: 0.16.0). For fully connected layers, torchao did not VOLUME ,
<Society logo(s) and publication title will appear here.>
TABLE 9: Parameters for complete and partitioned fullyconnected (FC) and matrix-matrix-multiplication (MM) layers implemented with OctopuScheduler (a) complete FC and MM layer parameters Layer Name
H
Input W
FC 1 FC 2
1 1
512 3072
MM 1
512
512
MM 2
512
3072
Weights H W
Output H W
512 3072
64 768
1 1
512
64
512
64
3072
768
512
768
64 768
(b) partitioned FC and MM layer parameters Layer Name
H
Input W
FC 1 FC 2
1 1
512 3072
MM 1
16
512
MM 2
16
3072
Weights H W
Output H W
Worker PEs
512 3072
32 16
1 1
32 16
2 48
512
16
16
16
128
3072
16
16
16
128
support a batch size of N ≤ 16 due to internal kernel limitations (see issue #2376), so the minimal batch size of N = 17 was selected. Similarly, OctopuScheduler runs fullyconnected layers always for batch size N = 4 m with m ∈ N due to hardware dimensions. On Jetson Orin Nano, the DNN layers were benchmarked using torch_tensorrt without any restrictions. G. E-prop Characterization
The code from [61] for training a recurrent spiking neural network for keyword spotting (Google Speech Commands [62]) with E-prop was ported to the SpiNNaker2 chip. Power and time measurements are shown in Table 10 and compared to the training with an NVIDIA V100 GPU. To cope with the low utilization of the SpiNNaker2 PEs (only 12 of 152 used), we calculate the effective power and effective energy which only consider the utilization fraction of the baseline power: Peffective = utilization × Pbaseline + Pdynamic Eeffective = Peffective × T,
TABLE 10: E-prop characterization Measurement
NVIDIA V100
SpiNNaker2
Batch size Utilization [%] Baseline Power [W] Dynamic Power [W] Effective Power [W] Time [h:m]
100 27 44 47.3 59.18 1:58
1 7.9 0.28558 0.0063 0.02886 500:00
Effective Energy [kJ]
419.0
52.0
(1) (2)
where T is the full training time. We note that the same method as in [61] is used that streams the input data from a host computer to a ping-pong buffer on SpiNNaker2. We expect that the training could be significantly accelerated by storing data in DRAM and by moving to data-parallel training as done for EventProp [50]. This, however, is out of scope of this paper.
VOLUME ,
19