arXiv:2604.15128v1 [cs.AR] 16 Apr 2026
SCENIC: Stream Computation-Enhanced SmartNIC Benjamin Ramhorst∗
Maximilian J. Heer∗
Luhao Liu
Heejae Kim
[email protected] ETH Zurich Zurich, Switzerland
[email protected] ETH Zurich Zurich, Switzerland
[email protected] ETH Zurich Zurich, Switzerland
[email protected] Seoul National University Seoul, Korea
Jonas Dann
Jin-Soo Kim
Gustavo Alonso
[email protected] ETH Zurich Zurich, Switzerland
[email protected] Seoul National University Seoul, Korea
[email protected] ETH Zurich Zurich, Switzerland
Abstract
focus has shifted to augment and optimize networking with compute through Smart Network Interface Cards (SmartNICs). SmartNICs, in addition to implementing the packet processing pipeline, also offload various steps of the data processing and compute pipeline. This includes, for example, network virtualization functions, storage access, network security and transport protocols, as shown by large-scale deployments of AWS Nitro [6], Microsoft AccelNet [26] and Alibaba CIPU [20]. Furthermore, off-the-shelf SmartNICs, such as NVIDIA Bluefield [60], AMD Pensando [11] and Broadcom Stingray [17], have become the backbone of modern ML systems, forming a scale-out network for thousands of GPUs [62]. In research, SmartNICs have been used to explore in-network compute for ML systems [30, 82], storage offloads [48, 84], databases [49, 53, 75], and security [63, 87]. Despite their high network bandwidth and ease-of-use, the very nature of closed-source and hardened commercial SmartNICs hinders novel research and adaptation to modern workloads. For example, next-generation protocols (e.g. UltraEthernet) [35, 70, 73] and congestion control algorithms [50, 85] are explored primarily through simulation due to the limited customizability of commercial NICs. Additionally, these SmartNICs often suffer from poor offload performance, as they are typically implemented using off-path Arm cores with high memory access latency [80] or on-path RISC-V cores with limited single-thread performance [19]. There has also been an increasing amount of interest in SmartNICs from the research community, with projects exploring many variations of the idea [12, 18, 27, 88]. However, most of these systems are limited in bandwidth (100G and often less), lack support for transport protocol offloading (e.g., RDMA, TCP/IP), have no native OS integration (e.g., through Linux netdev or ibv_device), and have limited or no integration with GPUs or SSDs like commercial SmartNICs do today. With some exceptions, they are not maintained, being just prototypes to demonstrate an idea or the potential to offload some functionality to NICs. In this paper, we present SCENIC, an FPGA-based SmartNIC with end-to-end system integration, designed to support in-network data processing and enabling full customization. SCENIC exploits the streaming nature of network traffic by
Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limited support for customization and data processing offload. In contrast, research SmartNICs often suffer from low bandwidth, limited functionality, and poor software compatibility - to the point that many are not actual NICs in a technical sense. This gap can be closed by treating the NIC datapath as a first-class stream computation substrate with shared hardware/software abstractions for a tight co-design of infrastructure and applications. To demonstrate this, we introduce SCENIC, an open-source datacenter SmartNIC. SCENIC implements a 200G network datapath over offloaded TCP/IP and RDMA stacks, as well as a fallback path for processing arbitrary network traffic. On top of the network logic, SCENIC combines on-datapath Stream Compute Units (SCUs) for data processing and embedded ARM cores for flexible control path manipulation with direct access to GPUs and SSDs. SCENIC is fully integrated with the OS, exposing native Linux network and RDMA verb interfaces, making the programmable datapath transparent to existing applications while enabling control of, e.g., user-defined offloads and programmable congestion control. SCENIC’s performance matches commercial platforms, and we show its versatility through several use cases such as offloaded collective communication and network-to-GPU hash-based data partitioning.
1
Introduction
With the ever-increasing scale of modern applications and improvements in compute efficiency through hardware specialization [1, 76], computer systems are becoming increasingly constrained by network performance [86]. Following current trends in compute and network bandwidth scaling, estimates indicate that distributed communication will make up between 50% and 75% of future training run-time [65]. At the same time, around 25-30% of CPU cycles in datacenters are spent on infrastructure tasks, often referred to as the "datacenter tax" [43]. To tackle these problems, recent ∗ Equal contribution.
1
introducing the notion of reprogrammable Stream CompuDPDK). SmartNICs extend conventional NICs by integrattation Units (SCU) that can be assigned to process network ing programmable compute on the card [4, 24, 25, 44, 77], flows in arbitrary ways. The SCUs can be utilized with any while also ensuring compatibility with existing networking of the offloaded network stacks (RDMA, TCP/IP) as well as frameworks. In the following, we discuss related platforms in combination with collective communication primitives. and compare them to SCENIC. We explicitly distinguish beSCENIC supports up to 200G bandwidth with native OS tween feature-complete SmartNICs and application-specific integration through Linux netdev and ibv_device. The renetworking platforms and components. sulting system is similar to those used at scale by Microsoft 2.1 Feature-complete SmartNICs in Azure cloud [69] or NVIDIA to manage communication for AI accelerators [13], but with SCENIC being an openThe central role of networking in cloud and datacenter comsource project providing higher customization possibilities puting has led to a number of academic SmartNICs based across the entire stack1 . SCENIC builds on well-established, on FPGAs (Table 1). Corundum [27] and OpenNIC [12] repopen-source projects (shell [45, 67], network stacks [28, 33], resent relatively early designs that focus on pure network communication libraries [32], and applications) that had to connectivity with added compute for control plane operbe redesigned for higher bandwidth, new FPGA architectures ations. Consequently, neither provides offloaded network and native datacenter compatibility. As such, it offers the stacks (RDMA or TCP/IP) nor integration with GPUs or possibility of modifying or replacing all of its components, SSDs. RecoNIC [88] extends OpenNIC with a hardwaretailoring the system for individual applications or research offloaded RDMA stack enabling direct, zero-copy data transuse cases. SCENIC’s contributions include: fer over RoCEv2. hXDP [18] focuses on offloading Linux XDP/eBPF programs on FPGAs hardware, representing a • A high-performance network datapath with offloads programming model orthogonal to SCENIC. It builds on top for common protocols (RDMA, TCP/IP), as well as a of NetFPGA [90], which provides a reference design for the fallback path for processing arbitrary traffic. network driver and the hardware implementation. Similar • Support for multiple parallel, isolated applications on to OpenNIC and Corundum, it does not include a complete the network datapath. Uniquely, SCENIC enables the deployment of offloaded applications both in programmable network stack nor integration with other devices (GPUs, SSDs). Additionally, all of the aforementioned platforms are logic, similar to FPGA-based SmartNICs, and in onbandwidth-bound to 100G or less. The same is true for Zechip Arm cores, similar to commercial DPUs. roNIC [74], which, similar to SCENIC, aims at datacenter • A virtual memory model which allows offloaded apcompatibility through support for GPUs and exposure as plications to access both NVIDIA and AMD GPUs, as netdev and ibv_device. ZeroNIC’s programmability lies in well as conventional SSDs. its flexible, software-defined control plane, rather than in • Native integration with Linux netdev and ibv_device, on-NIC data processing capabilities. exposing SCENIC as a standard NIC and ensuring comTurning to commercial platforms, a similar analysis can be patibility with existing networked applications. made. The Broadcom Stingray [17] targets general-purpose • A thorough evaluation covering throughput, latency, infrastructure offload with ARM cores and hardware accelGPU/SSD integration, and fairness, demonstrating pererators for crypto, RAID, and storage, but predates the curformance comparable to commercial 200G devices. rent generation of DPU platforms in both bandwidth and • A demonstration of SCENIC’s capabilities through two programmability. Intel IPU E2000 [40], NVIDIA BlueFielduse cases: SmartNIC-offloaded collective communica3 [60], and AMD Pensando Elba [11] combine hardware tion with performance comparable to OpenMPI, and RoCE engines, ARM SoC cores, IPSec and virtualization ofhash-based data partitioning for multi-GPU execution floads, storage acceleration and GPU-centric networking of database operators, with a 6.7x improvement over at up to 2×200G. All three provide programmable datapath the CPU baseline. capabilities through P4-programmable ARM cores (E2000), multi-threaded RISC-V cores (BF3), and P4-programmable 2 Related Work match-action pipelines (Pensando), which are generally more Throughout this paper, we adopt the definition of a NIC constrained than a fully customizable datapath on an FPGA. from [21]: a PCIe device with (i) a physical network interface The FPGA-based MangoBoost BoostX [56] explores a more (e.g., Ethernet or InfiniBand), (ii) a DMA engine for hostprogrammable variant at up to 400G with user-specified onmemory packet transfers, and (iii) an MSI-X interrupt interFPGA logic. However, MangoBoost’s own documentation face for completions. To preserve compatibility with decades states that it may contain forward-looking statements which of existing work and avoid software rewrites, we further reare subject to change, so some features reported in Table 1 quire a NIC to provide a host driver compatible with standard may not actually reflect the shipped product. All of the comnetworking frameworks (e.g., Linux netdev, ibv_device, or mercial platforms are closed-source, proprietary products 1 GitHub repository: https://github.com/fpgasystems/SCENIC that limit customization and research use [44]. In contrast, 2
netdev driver
ibv_device impl.
ARM Cores
Stream Compute
PCC
Open-Source
Direct-to-GPU
Direct-to-Storage
200 G
RDMA Offload
SCENIC (ours)
not supported.
TCP Offload
System
Network Speed
Table 1. Comparison of SmartNICs and related projects. ✓ supported; +□o partially supported;
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓ ✓ ✓ ✓ ✓
✓
Group 1: Academic FPGA-based NICs Corundum [27] OpenNIC [12] RecoNIC [88] hXDP [18] ZeroNIC [74]
100 G 100 G 100 G 40 G 100 G
+□o
✓ ✓ ✓
✓ ✓ +□o ✓ +□o
✓
✓ ✓ ✓ ✓ ✓
Group 2: Commercial NICs / DPUs Broadcom Stingray [17] Intel IPU E2000 [40] NVIDIA BlueField-3 B3220 [60] AMD Pensando Elba [11] Mango BoostX [56]
25/100 G 200 G 2x200 G
+□o +□o
✓ ✓ ✓
✓ ✓ ✓
✓ ✓ ✓
✓ ✓ ✓
+□o +□o
+□o ✓
✓ ✓ ✓
2x200 G ≤400 G
+□o ✓
✓ ✓
✓ ✓
✓ ✓
✓ ✓
+□o ✓
✓ ✓
✓ ✓
SCENIC provides comparable features and performance but as an open-source platform that combines offloaded network stacks, programmable congestion control, and GPU/SSDDirect coupled with native Linux integration — a combination that no existing open research platform achieves — while remaining vendor-agnostic and fully accessible to the research community. 2.2
processing pipelines. On the transport level, ACCL+ [32] implements collective communications on FPGAs, achieving performance comparable to software-based MPI, while FpgaNIC [79] proposes a networked FPGA platform with direct FPGA-to-GPU DMA. SuperNIC [51] addresses multitenancy on FPGA-based SmartNICs by introducing dynamically scheduled network task chains that share FPGA fabric resources across tenants. A challenge in the design of SCENIC was to combine many of these ideas and incorporate them into a single, efficient design.
Special-purpose research platforms
Besides feature-complete NICs, both industry and academia have explored network platforms that target specific aspects of the network rather than providing a general-purpose host interface. AccelTCP [57] accelerates TCP management by offloading connection states to ARM cores on a SmartNIC, bypassing the host kernel on the fast path. iPipe [54] further extends this approach and relocates distributed application logic directly onto the NIC using an actor-based programming model. On FPGAs, a number of projects explore offloaded networking stacks. FlexTOE [71], Limago [68] and EasyNet [31] propose full implementations of TCP/IP in FPGA fabric, while BALBOA [33] proposes an RDMA stack with a focus on datacenter compatibility. StRoM [72] extends an FPGA-based RoCEv2 stack with new opcodes for improved remote memory access in disaggregated memory systems. ClickNP [47] takes a more customizable approach, providing a modular programming framework for composing high-throughput packet processing pipelines, while FlowBlaze [66] focuses on hardware abstractions for stateful
3
Motivation & Design Requirements
Drawing on prior work and the characteristics of modern datacenters, we identify five requirements for our design: R1 – Performance: Networking in the cloud and datacenters is rapidly shifting from 100G to 200G/400G [29, 34]. Research SmartNICs, based on FPGAs, fail to achieve these bandwidths due to design complexity and clock frequency limitations. On the other hand, commercial SmartNICs achieve high bandwidth but rely on off-path ARM/RISCV cores, which are unsuitable for latency-sensitive, highthroughput offloads [19, 80]. SCENIC achieves the best of both worlds, achieving 200G with customizable, on-path streaming, and off-path ARM core offloads. R2 – Datacenter integration and compatibility: A SmartNIC requires well-defined software and driver interfaces and support for standard transport protocols. Beyond 3
SCENIC (FPGA) Arbitrary A4 packet processing
Off-path ARM cores Drivers
rdma-core F2
Linux netdev
E Example SCU: Flow monitoring
F1
TCP/IP
ibv_device SCU char device
Auxiliary devices
Congestion control
On-path reconfigurable SCUs
C
A3 c RoCE v2
Example SCU: Hash-based packet steering
D
System infrastructure Virtualization
B1
Arbitration
B2
PHY
G Python runtime
PCIe DMA
C++ run-time
MAC
Software
Traffic filter
SCENIC (CPU)
Novel transport stacks
Customizable network stacks
A2
A1
Figure 1. Overview of SCENIC with two example offloads: hash-based network-to-GPU data partitioning (Section 9.2) and hybrid flow monitoring (Section 6.2).
4
that, modern workloads require direct interaction with heterogeneous GPUs [81] and storage [15] at line rate. The challenge is doing so while also supporting in-network compute. SCENIC demonstrates both: compliance with existing software (exposed as Linux netdev and ibv_device) and direct interoperability with GPUs and SSDs. R3 – Customization and programmability: To utilize the available bandwidth, NIC customizability is key. As an example, DCQCN [89], the congestion control algorithm hardwired into today’s commodity RDMA NICs, is demonstrably suboptimal for modern traffic patterns [50, 85]. Yet, closedsource firmware and hardened hardware make it difficult to replace it or modify the transport layer to support novel semantics. SCENIC exposes the entire transport pipeline — from congestion control to the RDMA stack — as open-source FPGA IPs, enabling full customization and extensibility. R4 – Support for multiple flows and isolation: A production server hosts multiple tenants simultaneously, making isolation and fairness key requirements for a SmartNIC. However, many available SmartNICs often fall short of providing such guarantees due to the lack of virtualization and access control rules [44]. SCENIC addresses these through memory virtualization and flow arbitration, allowing the deployment of multiple, parallel and independent Stream Computation Units (SCUs). R5 – Programming model: Utilizing SmartNIC compute capabilities requires a user-exposed programming interface. In commercial SmartNICs, this typically mandates vendor lock-in via proprietary frameworks [44] such as DOCA [61] or Pensando SSDK [2]. Instead, SCENIC offers an open programming standard for offloads, supporting hardware description languages (HDLs), high-level synthesis (HLS), and network-specific languages such as P4.
System overview
SCENIC (Figure 1) consists of a modular hardware design, kernel-space drivers, and a user-space software API. In hardware, SCENIC implements a network datapath with IP cores for the MAC and PHY layers A1 , a traffic filter A2 , offloaded TCP/IP and RDMA stacks A3 , and a slow path for processing generic traffic A4 . Additional hardware components handle I/O virtualization B1 and fair resource sharing in multi-tenant systems B2 . Finally, the DMA engine C is used for host-to-SCENIC data movement and interrupts. The offloaded applications, termed Stream Computation Units (SCUs) D , can be used to implement custom functionality with access to all incoming and outgoing network traffic, as well as CPU/GPU memory and SSDs. On the driver-side, SCENIC exposes a full Linux netdev E and implements a rdma-core verbs provider through a kernel-space driver F1 and standard user-space libraries F2 , exposing SCENIC as an RDMA-capable SmartNIC with support for standard IB Verbs. SCENIC further extends Coyote’s driver stack with a host-side runtime G for offload configuration and control, including a Python runtime compatible with standard libraries such as NumPy and Pandas. SCENIC is designed with modularity in mind so that specific parts of the system can easily be enabled or disabled to produce application-specific designs. For example, workloads that do not require TCP/IP can disable it, freeing resources for SCUs. The only component always present is the slow path, ensuring Linux netdev compatibility. SCENIC’s design targets datacenter workloads and includes the slow path, an offloaded RDMA stack, memory virtualization for CPU and GPU access, and one SCU. TCP/IP, multiple SCUs2 ,
2 SCENIC can be configured to include up to 16 independent SCUs.
4
5
Hardware architecture
5.1
Network datapath
RDMA Commands
RDMA Completions Region #1
PCC
Region #0
Reconfiguration Signal
and access to SSD or FPGA memory (HBM/DDR) can be enabled via compile-time flags. We prototype SCENIC on a range of widely available FPGAs: 100G UltraScale+ AMD Alveo devices (U55C, U280, U250) and the most recent Alveo V80 (Versal architecture). The V80 provides substantially higher bandwidth (200G+), integrates Arm cores, and features a Network-on-Chip (NoC), making it a significantly more powerful platform. However, unlike UltraScale+ devices, the V80 has no complete shell, requiring all low-level hardware blocks (e.g., PCIe DMA, memory controllers, and reconfiguration logic) to be implemented from scratch. SCENIC bases its design on Coyote v2 [67], an open-source shell designed for 100G UltraScale+ platforms, and open-source TCP/IP [28] and RDMA [33] stacks. We redesigned many aspects of these projects to support higher throughput, direct GPU/SSD communication, and transparent software integration via netdev and ibv_device. Additionally, we configured the 200G AMD DCMAC IP [8] to enable networking on the V80, for which no fully functional reference design existed.
Flow Control
RoCE BALBOA InfiniBand-Hdrs.
DCQCN
QP-Info ECN RTT
UDP-Headers IP-Headers
ACK ECN
PHY
Figure 2. Programmable congestion control in SCENIC.
5.2
Congestion control and network extensions
Large-scale ML workloads expose the limitations of fixed congestion control algorithms [29, 50]. Programmable congestion control (PCC), with scenario-adaptive algorithm selection, addresses this challenge but requires direct modification of the transport logic, which is restricted on commercial NICs. SCENIC’s open architecture removes this constraint, enabling a broad design space for customization (R3). At the same time, any congestion control mechanism must satisfy the strict per-packet processing budget imposed by high link rates. At 200 G with MTU-sized RoCE packets, this budget is approximately:
To sustain line-rate performance (R1), SCENIC implements the physical layer on top of the Xilinx CMAC (UltraScale+ platforms, 100G) or DCMAC (Versal platforms, 200G) IP cores. These manage PAM2/PAM4 signal modulation [37, 39], forward error correction (FEC), and priority flow control (PFC) for lossless Ethernet [38]. Because these MAC cores stream data in a bandwidth-dependent clock domain and do not support backpressure, we safely cross clock domains into the application logic using a buffered, deeply pipelined datapath in both the RX and TX directions. Beyond the MAC, the hardware mirrors the OSI model. A networking prefilter acts as a triage layer, separating the fast path from the slow path: TCP and RoCEv2 packets are routed to our hardwareoffloaded networking stacks (if enabled), while all unhandled traffic is sent to the host via a dedicated DMA engine for Linux netdev processing (Section 7.1). For the offloaded fast path, an ARP resolver and a frame header insertion module handle the data link layer, exposing L2-stripped, network-layer packets on an AXI Stream bus. A subsequent L3/L4 filter then dispatches packets to the appropriate transport stack. SCENIC relies on the same opensource network stack IPs as Coyote [28, 33]. Crucially, these transport stacks decouple the control and data planes on the host-facing side. Extracted payloads are steered to specific Stream Computation Units (SCUs) based on control plane tags, such as the RoCE Queue Pair Number (QPN) or TCP Flow ID. This explicit, user-defined in-NIC routing enforces tenant and flow isolation (R4) and supports a fine-grained, per-packet programming model (R5).
4178 × 8 bit ≈ 167 ns 200 × 109 bit/s This constraint effectively rules out embedded CPU-based approaches. For example, NVIDIA BlueField-3 DOCA PCC cannot operate on a per-packet basis at line rate, and achieving such performance requires fundamental architectural changes demonstrated only by recent research [36]. In contrast, FPGA-based pipelines provide deterministic processing, handling each packet in a fixed number of clock cycles. At 391 MHz, 167 ns corresponds to roughly 65 cycles, which is sufficient even for algorithms that process multiple telemetry signals per packet, such as SMaRTT [16]. SCENIC leverages this property by implementing each congestion control algorithm as a dedicated hardware module, which can be swapped at runtime via dynamic partial reconfiguration of the FPGA fabric (R3). As reference implementations, SCENIC provides a simple ACK-clocked window-based flow controller and a full DCQCN implementation. In our experiments, we measure an average reconfiguration time of 8ms. This latency can be completely hidden through a dual-CC implementation as showcased in Figure 2. While one CCimplementation is actively steering the command flow, a second preloaded algorithm is already receiving congestion signals. When reconfiguration is triggered, this second module immediately takes over congestion control. 5
5.3
I/O virtualization and isolation
(e.g., the same interrupt would be launched for two packets arriving in a short amount of time). To prevent such cases, the SCENIC driver (Section 7.1) implements various locks and mutexes for shared resources, while also limiting the number of interrupt threads where applicable.
SCENIC extends Coyote’s virtual memory model, which enables the FPGA to access host CPU memory and, more recently, AMD GPU memory. However, for SmartNIC use cases, these are not sufficient. We therefore further extend the model to support NVIDIA GPUs as well as NVMe devices (R3). Access to NVIDIA GPU memory follows the same approach as for AMD GPUs, using the Linux dma-buf mechanism. This mechanism allows PCIe devices (e.g., GPUs) to export memory regions that can be imported by other PCIe devices (e.g., FPGAs), and has recently become the standard implementation for GPUDirect RDMA [59]. Beyond GPU memory, direct access to storage devices is equally critical for a SmartNIC that aims to minimize host CPU involvement on the I/O path. SCENIC therefore integrates a dedicated NVMe host controller within the FPGA fabric to handle both submission and completion logic. To bypass host CPU intervention, all NVMe control structures, including submission/completion queues and PRP lists, as well as data buffers are hosted in FPGA HBM/DDR and are mapped to a PCIe BAR. By maintaining direct access to NVMe doorbell registers and processing completion entries in hardware, SCENIC enables a low-latency data path between storage and network stacks. Coyote’s virtual memory model also ensures strict isolation (R4) between offloaded applications [67]. Resource fairness (R4) is achieved through system-wide arbiters, which ensure that all offloaded application equally share the available bandwidth through packet-based, round-robin arbitration. This enables SCENIC to natively support multiple RDMA QPs or TCP/IP sessions, each steered to a specific offload, with isolation and fairness guarantees. 5.4
6
Application offloads
6.1
Streaming on-datapath offloads
SCENIC’s SCUs implement custom user logic and can access host CPU, GPU, and SSD memory, as well as all incoming and outgoing network traffic. This enables streaming, dataflowstyle designs commonly used in network processing tasks such as compression, encryption, and hashing. What sets the SCUs apart from offloads on conventional SmartNICs is the combination of compute flexibility and memory bandwidth: FPGAs provide fine-grained control over parallelism, hardware primitives, and accesses patterns. While most SmartNICs rely on DRAM with tens of GB/s throughput, SCENIC supports HBM on platforms such as Alveo V80, U55C, and U280 with hundreds of GB/s of memory bandwidth. On other platforms (e.g., Alveo U250), it leverages on-board DRAM, while all supported FPGAs include tens of megabytes of on-chip Block- and Ultra-RAM (BRAM/URAM) with singlecycle access latency. This heterogeneous memory hierarchy, combined with fine-grained hardware control, enables highthroughput, latency-sensitive data processing (R1). Equally important is the SCU programming model (R5). As a network peripheral in the datacenter, SCENIC must allow SCUs to be developed quickly without requiring extensive hardware design knowledge. Thus, SCUs in SCENIC can be implemented in one of the following ways: • Register-transfer level (RTL), through languages such as SystemVerilog and VHDL. These languages give full control over the underlying hardware and best performance, but require considerable hardware expertise and long development cycles. • High Level Synthesis (HLS), a programming language based on C++ with pragmas to guide the hardware behavior. Developing with HLS is far more accessible, leading to up to 75% reduced development times, but often at the expense of resource consumption [46]. • SpinalHDL, a Scala-based hardware construction language that generates synthesizable RTL [64]. It provides higher abstraction than traditional RTL through parameterization, object-oriented, and functional programming, improving code reuse and productivity while preserving full hardware control. • P4, a language for packet processing pipelines, enabling rapid development of parsing, classification, and header manipulation. It is limited to match-action style processing and less suited for complex computations. Integration is done via AMD Vitis P4 IP [10], which synthesizes P4 code into FPGA RTL.
Host DMA engine
SCENIC implements a low-latency, high-throughput DMA engine for NIC-to-host transfers and interrupts (R1). On 100G UltraScale+ platforms, DMA is realized through the XDMA IP [7] supporting PCIe Gen3x16. On 200G Versal platforms, we use the hardened QDMA IP [9], which can be configured in PCIe Gen4x16 or Gen5x8 mode. Notably, the QDMA IP is designed for high-performance networking, with up to 2048 DMA queues. To maximize performance and reduce backpressure we configure the number of queues to match the number of outstanding packets. In both cases, like on commercial, ASIC-based NICs, we explicitly enable relaxed ordering for reads, reducing latency and preventing head-of-line transfer stalls. The same DMA engines are used to implement MSI-X interrupts between the SmartNIC and the host CPU, as required for netdev integration and other sources of interrupts. However, when processing interrupts, care has to be taken with parallelism: by default, multiple interrupt workers can be launched in parallel, which, while maximizing performance, can also lead to race conditions 6
6.2
HOST
SCENIC META PACKET
RX
napi_schedule
Interrupts
> irq_dispatch Register Settings PACKET
DMAcommand DMA
> net_poll
IRQ
META PACKET
PACKET DATA META
META Counting packet length
PHY
PACKET
MERGER
P4 SCUs are best suited for the slow path, which processes full headers, while HDL- and HLS-based SCUs are better suited for complex computational tasks where the header is stripped by the network stack (e.g., RDMA). In all cases, SCENIC’s build flow automatically compiles and synthesizes SCU code and links it with the rest of the platform. SCENIC also natively integrates with an open-source library of FPGA components3 , providing commonly used building blocks such as stream normalizers, hashing modules, crossbars, and data width converters. For AI workloads, SCENIC additionally integrates with ACCL+ [32], a collective offload engine for FPGAs with performance comparable to MPI. Finally, SCENIC incorporates an extensive simulation environment, enabling SCUs to be verified before deploying on hardware.
buff_vaddr buff_stride buff_size buff_tail irq_coal irq_time
TX PCIe
Figure 3. DMA packet forwarding to the network driver.
pipeline, classifying flows by source subnet to reflect podlevel positions in Fat Tree topologies [5], while implementation dynamic policy decision making on the Arm cores. A hardware timer periodically interrupts the CPU, which reads traffic statistics via the AXI bus. A dynamically configurable SCU rate limiter then enforces the resulting policies.
ARM-based off-datapath offloads
While existing literature points to the limitations of off-path Arm cores as primary means for data processing on SmartNICs [19, 80], SCENIC demonstrates how such cores can effectively complement the on-path SCUs for control plane operations. While dynamic SCU reconfiguration enables flexible data path adaptation, it also requires taking a flow offline during reconfiguration. Control plane updates, such as security rules or telemetry tasks, can instead be handled on an off-path core without interrupting the flow, satisfying the customizability requirement (R3). Additionally, the ability to run standard software on these cores provides a lightweight and accessible programming model (R5). On the V80 platform, SCENIC utilizes the embedded processing subsystem, consisting of a dual-core Arm® Cortex®A72 application processor and a dual-core Arm® Cortex®R5F real-time processor. Two types of interfaces enable tight coupling between the SCUs and the Arm cores: up to 16 IRQ connections can be dynamically assigned to individual SCUs providing low-latency event signaling from the data path to the Arm cores. Data transfer between the SCUs and the Arm cores is handled via a memory-mapped AXI bus, providing access to either control registers for lightweight command exchange or a dedicated memory unit for higher-capacity data buffering. The measured latency of this interface for a single access is about 0.3 𝜇s, and the single-trip time of an SCU interrupt reaching the handler in the Arm core is about 0.2 𝜇𝑠 — sufficient for transferring aggregated statistics, flow state snapshots, or control instructions to the processor cores without becoming a bottleneck for control plane operations. We demonstrate this SCU–CPU co-design with an incast traffic firewall. Stateful firewalls at 100G+ line rate exceed CPU capabilities, while FPGA-only solutions lack policy flexibility, requiring convoluted state machines [66]. SCENIC resolves this by offloading line-rate flow tracking to an SCU
7
Driver and software integration
In the following, we present SCENIC’s software API, as well as its native integration with Linux netdev and rdma-core to ensure compatibility with existing applications (R2). 7.1
Linux netdev integration
Figure 3 visualizes SCENIC’s arbitrary packet processing path which uses a hardware-software co-design to minimize DMA overhead and interrupt cost. On the RX path, a hardware state machine prepends a compact metadata tag, containing the packet length and a valid flag, directly before each packet payload. This allows the tag and the payload to share a single DMA transaction into a unified ring buffer, halving the number of required DMA operations compared to a naive two-transfer design and resulting in improved performance (R1). The driver’s net_poll handler then reads each metadata tag to determine packet validity and length before copying the payload starting right after the tag to the host networking stack. Interrupt delivery is managed by an MSI-X controller with two tunable behaviors: a timeout threshold guarantees bounded latency under sparse traffic, while interrupt coalescing amortizes interrupt overhead under high load, improving throughput. Both parameters are configurable by the driver, allowing the tradeoff to be tuned per workload. The TX path mirrors this structure of the RX path, using a ring buffer to enqueue outgoing commands. In terms of offloaded network acceleration, the blocks used for the physical network layer (CMAC and DCMAC) implement the Ethernet checksum calculation. Further features such as TSO or GRO are not implemented in hardware to keep the design small and resource-efficient, but can be added in future work.
3 https://github.com/fpgasystems/libstf
7
0.6 0.4 0.2 0.0
26
28 210 212 214 Msg. Size [B]
SCENIC (100G)
Hybrid
Throughput [Gbps]
Latency [ms]
0.8
10.0
based on Coyote’s software API. The software run-time allows various SCU-related control tasks, including setting and reading control registers, polling for interrupts, data movement and dynamic reconfiguration. Of particular interest is SCENIC’s Python run-time, which enables high-level control compatible with modern data processing libraries (e.g., PyTorch, Pandas). To ensure high performance, the Python run-time is implemented as a thin wrapper around the C++ run-time, thus minimizing overhead. An example of dynamically loading the application, setting control registers and checking completions is shown below.
7.5 5.0 2.5 0.0
Mellanox CX5 (100G)
Figure 4. Performance evaluation of the fallback path. Left: ping latency. Right: iperf3 throughput. Hybrid refers to Mellanox-to-SCENIC communication.
Code 1. Example Python code for SCU control # Load target application to SCU 2 ScenicReconfig().reconfigure_app( 2, "/path/to/scu/bitstream" )
7.2 ibv_device integration For data-intensive workloads, SCENIC includes a fully offloaded RoCEv2 stack and exposes itself as an ibv_device, ensuring compatibility with existing IB Verb applications (R2). This is realized through a two-component stack: (1) a kernel driver manages low-level hardware setup, including mapping of control registers and communication of link properties to the operating system, and (2) an ABI-defined interface connecting the driver to the scenic_ib provider implementation in the rdma-core userspace library. Accessing the mapped writeback and control registers allows the provider to interact directly with SCENIC’s DMA engine for outgoing RDMA WRITEs and RDMA READ REQUESTs and completion checking. Three design points illustrate the hardware-software codesign principle concretely. When calling ibv_reg_mr, the provider stores memory translations for the allocated userspace buffer directly in SCENIC’s on-device Translation Lookaside Buffers (TLBs). SCENIC’s driver ensures scalability across hundreds of QPs through automated translation lookups on TLB misses, while an LRU replacement strategy implemented in hardware preserves low latency for frequently active QPs (R1). Similarly, SCENIC’s completion mechanism avoids interrupt overhead entirely: the hardware performs atomic increments of per-QP completion counters, which the rdma-core provider polls directly, yielding high throughput without interrupt processing overhead. Finally, we implement ibv_create_qp_ex with an SCU index as an additional parameter, allowing applications to explicitly map RDMA flows to specific SCU processing pipelines (R3). Since each SCU maintains dedicated, non-shared datapath resources, this mapping provides hardware-level isolation between Queue Pairs assigned to different SCUs (R4). 7.3
# Create a thread and assign it to SCU 2 scenic_thread = ScenicThread(2) # Set control register, e.g., encryption key scenic_thread.set_csr(0x9f3c7a2b6e41d8c5, 0); # Check completions, for e.g., RDMA WRITE scenic_thread.get_completed("remote_write")
8
Evaluation
To evaluate SCENIC’s performance, we conduct microbenchmarks of the various components: slow-path packet processing, RDMA offload engine, GPU and SSD integration, and flow isolation. For a better understanding of the measurements, we provide comparisons with standard datacenter NICs, 100G Mellanox ConnectX-5 and 200G Broadcom. We evaluate SCENIC on the AMD Alveo U55C for 100G designs and on the AMD Alveo V80 for 200G designs. The measurements are conducted in a public research cluster, the AMDETH Heterogeneous Compute Cluster [58], which uses fully switched 100G and 200G networks. 8.1
Host networking performance
Due to the integration with netdev, we can use standard Linux tools to evaluate the performance of the slow path and the network driver for packets that cannot be processed by any of the offloaded networking stacks. To match the network offload found in SCENIC’s fallback path and measure the general datapath and driver performance, on-NIC segmentation offloads (TSO and GRO) are turned off on the commercial NICs. The average latency and jitter between SCENIC and commercial NICs through the ping utility (ICMP request and response) over a 100G link is shown in Figure 4. Commercial, ASIC-based NICs likely benefit from generally higher clock speeds and further optimized DMA
Software API
To control and configure the SCUs, SCENIC implements a high-level (R5) software run-time in both Python and C++, 8
Throughput [Gbps] Latency [µs]
200 150
200 READ
150
100
100
50
50
0 30 25 20 15 10 5 0
26 27 28 29 210 211 212 213 214 215 216 217 218 READ
26 27 28 29 210 211 212 213 214 215 216 217 218
0 30 25 20 15 10 5 0
WRITE
26 27 28 29 210 211 212 213 214 215 216 217 218 WRITE
26 27 28 29 210 211 212 213 214 215 216 217 218
Msg. Size [B] Mellanox CX5 (100G)
Msg. Size [B] Broadcom (200G)
SCENIC (100G)
SCENIC (200G)
Figure 5. RDMA performance benchmark in a fully switched datacenter network. engines, resulting in slightly lower latencies. Despite this gap, SCENIC’s slow-path latency is well within the range required for management traffic: SSH sessions, monitoring, and control-plane communication remain fully responsive under all tested conditions. A similar conclusion can be drawn for achievable throughput with iperf3 at a typical MTU of 1500B and no further optimizations. While falling slightly behind commercial NICs, the performance is sufficient for practical tasks like video streaming or scp-based data transfer. This is expected and acceptable by design since bulk data transfers are handled at line rate by the networking stacks. 8.2
and polls for local delivery; for WRITE benchmarks, the requester writes data to the server’s GPU memory and polls on acknowledgments generated by the server-side NIC. Important to note, NIC-side acknowledgments are generated independently from the GPU memory controller; however, we confirm similar throughput by also polling on memory contents of the server-side GPU. As shown in Figure 6, performance on AMD GPUs saturates the link and is comparable to our CPU baseline (Figure 5). In both cases, we observe lower performance on NVIDIA GPUs. Part of this can be attributed to a sub-optimal PCIe configuration of the NVIDIA node, over which we have limited control. We confirm the sub-optimal performance by measuring local (CPU-to-SCENIC) transfers and observe saturation at 20 GBps, rather than the expected 23 - 25 GBps for a PCIe Gen4x16 set-up. Additionally, READs exhibit higher
RDMA performance
8.3
Throughput [Gbps]
SCENIC’s full exposure as ibv_device enables the standard perftest benchmark [52] to be run without modifications, providing a direct comparison to commercial NICs (Figure 5). Latency measurements via ib_write_lat and ib_read_lat show a slight advantage for the Mellanox, consistent with a more mature PCIe DMA engine and the higher clock speed of the ASIC NIC. However, this gap is modest and does not impact SCENIC’s target use cases, for which throughput and programmability represent the main objectives. For throughput, SCENIC saturates the available network bandwidth for both ib_write_bw and ib_read_bw and achieves performance comparable to commercial NICs.
200 150 100 50 0
26 27 28 29 210 211 212 213 214 215 216 217 218 Msg. Size [B]
GPU integration AMD RD
We measure end-to-end throughput between SCENIC and GPUs using RDMA READs and WRITEs. In this configuration, the GPU node acts as the responder (server). For READ benchmarks, the client reads data from GPU memory
AMD WR
Nvidia RD
Nvidia WR
Figure 6. Throughput of SCENIC to GPU with RDMA READs and RDMA WRITEs. 9
20
150
15
100
10
50
5 0
4
8
16 32 64 Chunk Size [KiB]
SCENIC Throughput Host Throughput
128
Throughput [Gbps]
200 Latency [µs]
Throughput [Gbps]
25
0
SCENIC Latency Host Latency
100 50 0
Flow #0 Flow #1 Flow #2 Flow #3
1 Flow
2 Flows
3 Flows
4 Flows
To demonstrate this design aspect, we configure SCENIC with four SCUs and run a throughput benchmark with 128 KiB RDMA READs. We incrementally scale the workload from one to four parallel flows, mapping each distinct flow to its own SCU. Figure 8 shows a time series of the throughput as more flows are added to the system. As required, the aggregate throughput saturates the expected READ-bandwidth over a 200G link while being equally shared among active flows, even with new ones being added. This demonstrates that the SCU-based architecture prevents cross-flow interference, ensuring that independent network streams do not contend for bandwidth-constrained resources under full load.
performance degradation on NVIDIA GPUs and warrant further investigation. However, prior works, on both AMD [78] and NVIDIA [74] GPUs have noted lower performance for GPU-FPGA transfers, especially for READs. We plan on addressing these bottlenecks in future, but the presented results showcase SCENIC’s capability for vendor-agnostic integration with GPUs.
8.6
SSD integration over TCP/IP
Resource consumption
The resource consumption of the core SCENIC design, consisting of the slow path, an offloaded RDMA stack, the memory virtualization layer, one SCU and the host DMA engine is shown in Table 2. SCENIC’s resource consumption remains low: less than 30% on the U55C/U250 and less than 15% on the V80, leaving ample room for offloaded applications.
We evaluate SCENIC’s TCP/IP offload engine on a networkto-storage path. On the U55C platform, we construct a pipeline in which incoming TCP/IP segments are reassembled in hardware and written directly into a local NVMe SSD, fully bypassing the host CPU. We compare against a host baseline using a 100G Mellanox ConnectX-5, where incoming data is received through the kernel TCP stack and written to the SSD via a single-threaded io_uring event loop with O_DIRECT. Figure 7 compares the latency and throughput of both configurations. SCENIC achieves 2–3× lower per-completion latency across all chunk sizes (e.g., 25.6 𝜇s vs. 72.4 𝜇s at 4 KB). A breakdown of the end-to-end latency reveals that software overheads, including kernel TCP/IP packet processing, system call transitions, and buffer management, dominate the host path (56–115 𝜇s), whereas SCENIC’s hardware path only adds 14–27 𝜇s. In terms of throughput, even with 8 outstanding requests, SCENIC saturates the NVMe bandwidth and outperforms the host’s best configuration with 64 outstanding requests by 1.27–1.47× across all chunk sizes. Moreover, since SCENIC entirely bypasses the host CPU, the architecture is expected to scale to multiple SSDs without the software overhead. 8.5
150
Figure 8. Time series of bandwidth sharing scaling up to four parallel flows performing 128 KiB RDMA READs through separated SCUs.
Figure 7. TCP to NVMe throughput and latency with SCENIC offload, compared to the host-side software baseline.
8.4
200
9
Use cases
9.1
ACCL
As a first use case, we deploy the open-source ACCL+ [32, 83] library on SCENIC, enabling offloaded collective communication. Figure 9 shows the results of a four-node benchmark comparing the BROADCAST and GATHER collectives on SCENIC with ACCL+ against the CPU baseline with RDMA OpenMPI and a Mellanox Connect-X5. SCENIC matches the performance of the commercial NIC with established libraries, while providing two clear advantages. First, collective communication represents a large part of the datacenter tax and the network bottleneck in large-scale AI Table 2. SCENIC resource consumption. Platform V80 U55C U250
Performance isolation through separate SCUs
The strict, hardware-level separation of network flows through different SCUs does not only allow fine-grained, per-flow processing, but also guarantees fairness through arbitration (R4). 10
LUT [%] 11.5 28.8 22.5
REG [%] 12.5 25.8 19.7
BRAM [%] 17.1 25.2 18.8
Latency [µs]
Gather
103
102
Throughput [Gbps]
Broadcast
103
102
101 2
13
15
17
19
2 2 2 2 Msg. Size [B]
21
OpenMPI
101
2
13
15
17
19
2 2 2 2 Msg. Size [B]
80 60 40 20 0
21
215
213
217
219
221
Number of Rows B(1)
SCENIC
B(16)
SCENIC
(b) Latency Latency [ms]
Figure 9. Comparison of BROADCAST and GATHER collectives on SCENIC with OpenMPI on a commercial NIC.
training [41, 55]; offloading them to the network can free up CPU cycles or reduce the GPU utilization [14]. Second, offloaded collectives create the possibility of collocating gradient compression as an in-network processing step to further overlap compute and communication [3]. For future work we plan to extend this example into an end-to-end pipeline for large-scale model training on GPUs with SmartNIC-offloaded collectives and compression. 9.2
(a) Throughput (threads)
100
10 6.7× 5 0
218
219
220
221
Number of Rows B(1): Partitioning B(1): GPU comm.
B(1): RDMA comm. SCENIC
Figure 10. Performance of hash partitioning on the CPU (B: Baseline, 1 and 16 threads) and offloaded with SCENIC.
Hash-based data partitioning
As a second use case, we demonstrate SCENIC as an accelerator for cloud-native data processing systems. SmartNICs have emerged as a compelling platform for line-rate data decoding and parsing [23], as well as for offloading data scanning and filtering [22]—operations that impose significant runtime and memory overheads on host CPUs. With this application, we target hash partitioning, a critical building block for multi-GPU query execution [42] in modern data processing systems. Hash partitioning is used to deterministically split up data into equal parts that can be processed independently by, e.g., a join or aggregation operator. We implement hash partitioning as a SCENIC SCU. The SCU maintains an on-chip hash buffer (16 × 216 hashes in our configuration) which supports hash folding over composite key columns. The resulting hashes are then used to partition a set of data columns. To support data sets that exceed the buffer capacity (> 219 rows), we use batching. For each GPU in the system, a dedicated pipeline selects the payload rows for each data columns and accumulates them in an output buffer, which is flushed in transfers of 64kB (the smallest transfer size that saturates PCIe throughput). We evaluate this use case on a two-node setup: a data processing node (AMD U55C and 4 AMD MI210 GPUs) and a remote memory node (AMD U55C). We generate a synthetic two-column table—one key column for hashing and one data column as the payload that is distributed across GPUs. Figure 10 shows the throughput and the latency for
the multi-threaded software baseline compared to SCENICoffloaded hash partitioning. SCENIC offloading achieves latency that scales linearly with data set size and, for larger transfer sizes, approaches the lower bound of just the RDMA communication. Throughput shows a fixed overhead at small data set sizes that is amortized for larger data sets. Performance drops slightly beyond the on-chip hash buffer capacity, where batching is required. The software baseline incurs substantially higher latency and reaches lower throughput, with thread management overhead limiting scalability at small transfer sizes. This use case demonstrates one of the many possibilities offered by SCENIC to accelerate applications and move data to where it can be best executed, thereby improving the utilization of expensive accelerators.
10
Conclusions and future work
In this paper, we presented SCENIC, a fully open-source, 200G, FPGA-based SmartNIC suitable for deployment in heterogeneous compute environments with CPUs, GPUs and SSDs. SCENIC demonstrates several innovative aspects in FPGA design for scalability, compatibility across architectures, and flexibility in network support for different protocols and features. It does this while being a full featured NIC rather than just focused on a single function, as most research prototypes do. Due to its native integration with Linux netdev and rdma-core, SCENIC works out of the 11
box with existing applications, while providing easy-to-use Python and C++ run-times for offload configuration and control. Reconfigurable SCUs and off-datapath Arm cores enable low-latency, line-rate dataflow offloads, as shown by hash-based network-to-GPU data partitioning, achieving a 6.7× latency improvement over the baseline. In the future, we plan to further improve SCENIC in a variety of ways, with a primary emphasis on performance scaling to 2x200G with bifurcated PCIe Gen5x16, leading to up to 400G throughput. Additionally, we plan to explore hybrid application offloads, leveraging both the off-path ARM cores and the SCUs with suitable workload partitioning between the two. Finally, future work will involve new network protocols (e.g., UltraEthernet) and congestion control algorithms with SCENIC as a starting point.
[6] Amazon Web Services. 2022. The Components of the Nitro System (The Security Design of the AWS Nitro System Whitepaper). Technical Report. Amazon Web Services. https://docs.aws.amazon.com/whitepapers/la test/security-design-of-aws-nitro-system/the-components-of-thenitro-system.html Accessed: 2026-04-15. [7] AMD. 2025. DMA/Bridge Subsystem for PCI Express Product Guide (PG195). https://docs.amd.com/r/en-US/pg195-pcie-dma [8] AMD. 2025. Versal Adaptive SoC 600G Channelized Multirate Ethernet Subsystem (DCMAC) LogiCORE IP Product Guide (PG369). https: //docs.amd.com/r/en-US/pg369-dcmac/Introduction [9] AMD. 2025. Versal Adaptive SoC CPM DMA and Bridge Mode for PCI Express v3.4. https://docs.amd.com/r/en-US/pg347-cpm-dmabridge?tocId=oTd_ZrdYcOWw7fqmc3hb9g [10] AMD. 2025. Vitis Networking P4. https://docs.amd.com/r/enUS/ug1308-vitis-p4-user-guide [11] AMD Pensando. 2022. AMD Pensando Elba DPU (DSC-200) Product Overview. https://www.amd.com/en/products/data-processingunits/pensando.html. [12] AMD/Xilinx. 2021. OpenNIC: An Open-Source NIC Shell for Alveo FPGAs. GitHub. https://github.com/Xilinx/open-nic. [13] Kyle Aubrey and Farshad Ghodsian. 2026. Inside NVIDIA Groq 3 LPX: The Low-Latency Inference Accelerator for the NVIDIA Vera Rubin Platform. NVIDIA Technical Blog. https://developer.nvidia.com/blo g/inside-nvidia-groq-3-lpx-the-low-latency-inference-acceleratorfor-the-nvidia-vera-rubin-platform/ Accessed: 2026-03-28. [14] John Bachan, Kaiming Ouyang, Misbah Mubarak, Thomas Gillis, Bruce Chang, Devendar Bureddy, Giuseppe Congiu, Keith Caton, Kyle Aubrey, and Xiaofan Li. 2025. Enabling Fast Inference and Resilient Training with NCCL 2.27. https://developer.nvidia.com/blog/enablingfast-inference-and-resilient-training-with-nccl-2-27/ [15] Wei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre, Paramvir Bahl, Ameya Bhagat, Gowri Bhaskara, Tanya Brokhman, Lei Cao, Ahmad Cheema, Rebecca Chow, Jeff Cohen, Mahmoud Elhaddad, Vivek Ette, Igal Figlin, Daniel Firestone, Mathew George, Ilya German, Lakhmeet Ghai, Eric Green, Albert G. Greenberg, Manish Gupta, Randy Haagens, Matthew Hendel, Ridwan Howlader, Neetha John, Julia Johnstone, Tom Jolly, Greg Kramer, David Kruse, Ankit Kumar, Erica Lan, Ivan Lee, Avi Levy, Marina Lipshteyn, Xin Liu, Chen Liu, Guohan Lu, Yuemin Lu, Xiakun Lu, Vadim Makhervaks, Ulad Malashanka, David A. Maltz, Ilias Marinos, Rohan Mehta, Sharda Murthi, Anup Namdhari, Aaron Ogus, Jitendra Padhye, Madhav Pandya, Douglas Phillips, Adrian Power, Suraj Puri, Shachar Raindel, Jordan Rhee, Anthony Russo, Maneesh Sah, Ali Sheriff, Chris Sparacino, Ashutosh Srivastava, Weixiang Sun, Nick Swanson, Fuhou Tian, Lukasz Tomczyk, Vamsi Vadlamuri, Alec Wolman, Ying Xie, Joyce Yom, Lihua Yuan, Yanzhao Zhang, and Brian Zill. 2023. Empowering Azure Storage with RDMA. In 20th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2023, Boston, MA, April 17-19, 2023, Mahesh Balakrishnan and Manya Ghobadi (Eds.). USENIX Association, 49–67. https://www.usenix.org/conference/nsdi23/presentation/bai [16] Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Anup Agarwal, Daniele De Sensi, Rong Pan, Costin Raiciu, Mark Handley, Mihai Brodschi, Timo Schneider, Nils Blach, Daniel Santos Ferreira Alves, and Torsten Hoefler. 2026. SMaRTT: Sender-based Marked Rapidlyadapting Trimmed & Timed Transport. (2026). arXiv:2404.01630 [cs.NI] https://arxiv.org/abs/2404.01630 [17] Broadcom. 2019. Broadcom Stingray PS225 Dual-Port 25GbE PCIe Ethernet SmartNIC Data Sheet. https://www.broadcom.com/compa ny/news/product-releases/53106. [18] Marco Spaziani Brunella, Giacomo Belocchi, Marco Bonola, Salvatore Pontarelli, Giuseppe Siracusano, Giuseppe Bianchi, Aniello Cammarano, Alessandro Palumbo, Luca Petrucci, and Roberto Bifulco. 2020. hXDP: Efficient Software Packet Processing on FPGA NICs. In 14th USENIX Symposium on Operating Systems Design and Implementation,
Acknowledgments We would like to thank AMD for the donation of the Heterogeneous Accelerated Compute Cluster (HACC) at ETHZ which was used for the development of this project and Geert Roks for the support and help with the cluster set-up. This work was funded in part through an unrestricted grant from AMD. Additionally, we would like to thank contributors of the SLASH project from AMD Research Dublin, in particular Alexandru Ulmămei, Lucian Petrica and Mario Ruiz, for insightful guidance and help with the V80 FPGA. We have used Google Gemini and Anthropic Claude for code debugging, language checking, as well as for the homogenization of the plotting scripts used to generate the diagrams.
References [1] Daniel Abadi, Anastasia Ailamaki, David G. Andersen, Peter Bailis, Magdalena Balazinska, Philip A. Bernstein, Peter Boncz, Surajit Chaudhuri, Alvin Cheung, AnHai Doan, Luna Dong, Michael J. Franklin, Juliana Freire, Alon Y. Halevy, Joseph M. Hellerstein, Stratos Idreos, Donald Kossmann, Tim Kraska, Sailesh Krishnamurthy, Volker Markl, Sergey Melnik, Tova Milo, C. Mohan, Thomas Neumann, Beng Chin Ooi, Fatma Ozcan, Jignesh M. Patel, Andrew Pavlo, Raluca A. Popa, Raghu Ramakrishnan, Christopher Ré, Michael Stonebraker, and Dan Suciu. 2022. The Seattle report on database research. Commun. ACM 65, 8 (2022), 72–79. doi:10.1145/3524284 [2] Advanced Micro Devices, Inc. 2024. AMD Pensando Software-in-Silicon Development Kit (SSDK). https://www.amd.com/content/dam/amd/ en/documents/pensando-technical-docs/product-briefs/pensandossdk-product-brief.pdf [3] Saurabh Agarwal, Hongyi Wang, Shivaram Venkataraman, and Dimitris S. Papailiopoulos. 2022. On the Utility of Gradient Compression in Distributed Training Systems. (2022). https://proceedings.mlsys.or g/paper_files/paper/2022/hash/773862fcc2e29f650d68960ba5bd1101Abstract.html [4] Olasupo Ajayi and Ryan Grant. 2025. A Chronological Analysis of the Evolution of SmartNICs. CoRR abs/2512.04054 (2025). arXiv:2512.04054 doi:10.48550/ARXIV.2512.04054 [5] Mohammad Al-Fares, Alexander Loukissas, and Amin Vahdat. 2008. A scalable, commodity data center network architecture. In Proceedings of the ACM SIGCOMM 2008 Conference on Data Communication (Seattle, WA, USA) (SIGCOMM ’08). Association for Computing Machinery, New York, NY, USA, 63–74. doi:10.1145/1402958.1402967 12
OSDI 2020, Virtual Event, November 4-6, 2020. USENIX Association, 973– 990. https://www.usenix.org/conference/osdi20/presentation/brunella [19] Xuzheng Chen, Jie Zhang, Ting Fu, Yifan Shen, Shu Ma, Kun Qian, Lingjun Zhu, Chao Shi, Yin Zhang, Ming Liu, and Zeke Wang. 2024. Demystifying Datapath Accelerator Enhanced Off-path SmartNIC. In 32nd IEEE International Conference on Network Protocols, ICNP 2024, Charleroi, Belgium, October 28-31, 2024. IEEE, 1–12. doi:10.1109/ICNP 61940.2024.10858560 [20] Alibaba Cloud Community. 2022. A Detailed Explanation about Alibaba Cloud CIPU. https://www.alibabacloud.com/blog/a-detailedexplanation-about-alibaba-cloud-cipu_599183 [21] Dan Daly, Jakub Kicinski, and Willem de Bruijn. 2023. OCP NIC Core Features Specification, Version 1.0. Technical Specification. Open Compute Project (OCP). https://www.opencompute.org/document s/ocp-server-nic-core-features-specification-ocp-spec-format-1-1pdf Accessed: 2026-03-23. [22] Jonas Dann and Gustavo Alonso. 2026. Should I Hide My Duck in the Lake? CoRR abs/2602.18775 (2026). doi:10.48550/ARXIV.2602.18775 [23] Jonas Dann, Royden Wagner, Daniel Ritter, Christian Faerber, and Holger Fröning. 2022. PipeJSON: Parsing JSON at Line Speed on FPGAs. In International Conference on Management of Data, DaMoN 2022, Philadelphia, PA, USA, 13 June 2022, Spyros Blanas and Norman May (Eds.). ACM, 3:1–3:7. doi:10.1145/3533737.3535094 [24] Tristan Döring, Henning Stubbe, and Kilian Holzinger. 2021. SmartNICs: Current Trends in Research and Industry. Technical Report NET2021-05-1. Chair of Network Architectures and Services, Department of Informatics, Technical University of Munich. https://www.net.in.t um.de/fileadmin/TUM/NET/NET-2021-05-1/NET-2021-05-1_05.pdf [25] Sergio Elizalde, Ali AlSabeh, Ali Mazloum, Samia Choueiri, Elie F. Kfoury, Jose Gomez, and Jorge Crichigno. 2025. A survey on security applications with SmartNICs: Taxonomy, implementations, challenges, and future trends. J. Netw. Comput. Appl. 242 (2025), 104257. doi:10.1 016/J.JNCA.2025.104257 [26] Daniel Firestone, Andrew Putnam, Sambrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian M. Caulfield, Eric S. Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitu Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, and Albert G. Greenberg. 2018. Azure Accelerated Networking: SmartNICs in the Public Cloud. In 15th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2018, Renton, WA, USA, April 9-11, 2018, Sujata Banerjee and Srinivasan Seshan (Eds.). USENIX Association, 51–66. https://www.usenix.org/conference/nsdi18/presentation/firestone [27] Alex Forencich, Alex C. Snoeren, George Porter, and George Papen. 2020. Corundum: An Open-Source 100-Gbps Nic. In 28th IEEE Annual International Symposium on Field-Programmable Custom Computing Machines, FCCM 2020, Fayetteville, AR, USA, May 3-6, 2020. IEEE, 38–46. doi:10.1109/FCCM48280.2020.00015 [28] fpgasystems. [n. d.]. GitHub - fpgasystems/fpga-network-stack: Scalable Network Stack for FPGAs (TCP/IP, RoCEv2). https://github.com /fpgasystems/fpga-network-stack [29] Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, Shuqiang Zhang, Mikel Jimenez Fernandez, Shashidhar Gandham, and Hongyi Zeng. 2024. RDMA over Ethernet for Distributed Training at Meta Scale. In Proceedings of the ACM SIGCOMM 2024 Conference, ACM SIGCOMM 2024, Sydney, NSW, Australia, August 4-8, 2024. ACM, 57–70. doi:10.1145/3651890.3672233 [30] Anqi Guo, Yuchen Hao, Xiteng Yao, Shining Yang, Jianyu Huang, Tony (Tong) Geng, and Martin Herbordt. 2025. SmartNIC-GPUCPU Heterogeneous System for Large Machine Learning Model with
Software-Hardware Codesign. In Proceedings of the 39th ACM International Conference on Supercomputing (ICS ’25). Association for Computing Machinery, New York, NY, USA, 837–852. doi:10.1145/3721145. 3729514 [31] Zhenhao He, Dario Korolija, and Gustavo Alonso. 2021. EasyNet: 100 Gbps Network for HLS. In 31st International Conference on FieldProgrammable Logic and Applications, FPL 2021, Dresden, Germany, August 30 - Sept. 3, 2021. IEEE, 197–203. doi:10.1109/FPL53798.2021.00 040 [32] Zhenhao He, Dario Korolija, Yu Zhu, Benjamin Ramhorst, Tristan Laan, Lucian Petrica, Michaela Blott, and Gustavo Alonso. 2024. ACCL+: an FPGA-Based Collective Engine for Distributed Applications. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, Ada Gavrilovska and Douglas B. Terry (Eds.). USENIX Association, 211–231. https: //www.usenix.org/conference/osdi24/presentation/he [33] Maximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu, Zhiyi Hu, Jonas Dann, and Gustavo Alonso. 2025. RoCE BALBOA: Serviceenhanced Data Center RDMA for SmartNICs. arXiv:2507.20412 doi:10 .48550/ARXIV.2507.20412 [34] Torsten Hoefler, Duncan Roweth, Keith D. Underwood, Robert Alverson, Mark Griswold, Vahid Tabatabaee, Mohan Kalkunte, Surendra Anubolu, Siyuan Shen, Moray McLaren, Abdul Kabbani, and Steve Scott. 2023. Data Center Ethernet and Remote Direct Memory Access: Issues at Hyperscale. Computer 56, 7 (2023), 67–77. doi:10.1109/MC.2023.3261184 [35] Torsten Hoefler, Karen Schramm, Eric Spada, Keith D. Underwood, Cedell Alexander, Bob Alverson, Paul Bottorff, Adrian M. Caulfield, Mark Handley, Cathy Huang, Costin Raiciu, Abdul Kabbani, Eugene Opsasnick, Rong Pan, Adee Ran, and Rip Sohan. 2025. Ultra Ethernet’s Design Principles and Architectural Innovations. arXiv:2508.08906 doi:10.48550/ARXIV.2508.08906 [36] Hongjing Huang, Jie Zhang, Xuzheng Chen, Ziyu Song, Jiajun Qin, and Zeke Wang. 2025. SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA Engine. In Proceedings of the 2025 USENIX Annual Technical Conference, USENIX ATC 2025, Boston, MA, USA, July 7-9, 2025, Deniz Altinbüken and Ryan Stutsman (Eds.). USENIX Association, 1243–1260. https://www.usenix.org/conference/atc25/presen tation/huang-hongjing [37] IEEE. 2010. IEEE Standard for Information technology–Local and metropolitan area networks–Specific requirements–Part 3: CSMA/CD Access Method and Physical Layer Specifications Amendment 4: Media Access Control Parameters, Physical Layers, and Management Parameters for 40 Gb/s and 100 Gb/s Operation. 457 pages. doi:10.1109/IEEESTD.2010.5501740 [38] IEEE. 2011. IEEE Standard for Local and metropolitan area networks– Media Access Control (MAC) Bridges and Virtual Bridged Local Area Networks–Amendment 17: Priority-based Flow Control. 40 pages. doi:10.1109/IEEESTD.2011.6032693 [39] IEEE. 2017. IEEE Standard for Ethernet - Amendment 10: Media Access Control Parameters, Physical Layers, and Management Parameters for 200 Gb/s and 400 Gb/s Operation. 416 pages. doi:10.1109/IEEESTD.20 17.8207825 [40] Intel. 2022. Intel Infrastructure Processing Unit (Intel IPU) E2000. https://www.intel.com/content/www/us/en/products/details/netwo rk-io/ipu.html. [41] Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi Zou, Sida Zhao, Liang Xiang, Zherui Liu, Zhe Li, Xiaoying Jia, Jianxi Ye, Xin Jin, and Xin Liu. 2024. MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs. In 21st USENIX Symposium on Networked Systems Design and Implementation, 13
NSDI 2024, Santa Clara, CA, April 15-17, 2024, Laurent Vanbever and Irene Zhang (Eds.). USENIX Association, 745–760. https://www.usen ix.org/conference/nsdi24/presentation/jiang-ziheng [42] Marko Kabic, Bowen Wu, Jonas Dann, and Gustavo Alonso. 2025. Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs. Proc. VLDB Endow. 18, 11 (2025), 4350–4363. doi:10 .14778/3749646.3749698 [43] Svilen Kanev, Juan Pablo Darago, Kim Hazelwood, Parthasarathy Ranganathan, Tipp Moseley, Gu-Yeon Wei, and David Brooks. 2015. Profiling a warehouse-scale computer. In Proceedings of the 42nd Annual International Symposium on Computer Architecture (Portland, Oregon) (ISCA ’15). Association for Computing Machinery, New York, NY, USA, 158–169. doi:10.1145/2749469.2750392 [44] Elie F. Kfoury, Samia Choueiri, Ali Mazloum, Ali AlSabeh, Jose Gomez, and Jorge Crichigno. 2024. A Comprehensive Survey on SmartNICs: Architectures, Development Models, Applications, and Research Directions. IEEE Access 12 (2024), 107297–107336. doi:10.1109/ACCESS.2 024.3437203 [45] Dario Korolija, Timothy Roscoe, and Gustavo Alonso. 2020. Do OS abstractions make sense on FPGAs?. In 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020, Virtual Event, November 4-6, 2020. USENIX Association, 991–1010. https://www.us enix.org/conference/osdi20/presentation/roscoe [46] Sakari Lahti and Timo D. Hämäläinen. 2025. High-Level Synthesis for FPGAs - A Hardware Engineer’s Perspective. IEEE Access 13 (2025), 28574–28593. doi:10.1109/ACCESS.2025.3540320 [47] Bojie Li, Kun Tan, Layong Larry Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, and Peng Cheng. 2016. ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable Hardware. In Proceedings of the ACM SIGCOMM 2016 Conference, Florianopolis, Brazil, August 22-26, 2016, Marinho P. Barcellos, Jon Crowcroft, Amin Vahdat, and Sachin Katti (Eds.). ACM, 1–14. doi:10.1145/2934872.2934897 [48] Jiayong Li, Jonas Dann, Zhenhao He, Gustavo Alonso, Sai Rahul Chalamalasetti, Dejan Milojicic, Lance Evans, Alex Veprinsky, and Runbin Shi. 2026. StreamDedup: Distributed In-line Deduplication for Disaggregated Storage. ACM Trans. Reconfigurable Technol. Syst. (March 2026). doi:10.1145/3799896 [49] Junru Li, Youyou Lu, Qing Wang, Jiazhen Lin, Zhe Yang, and Jiwu Shu. 2022. AlNiCo: SmartNIC-accelerated Contention-aware Request Scheduling for Transaction Processing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, 951– 966. https://www.usenix.org/conference/atc22/presentation/li-junru [50] Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang, Zheng Cao, Ming Zhang, Frank Kelly, Mohammad Alizadeh, and Minlan Yu. 2019. HPCC: high precision congestion control. In Proceedings of the ACM Special Interest Group on Data Communication, SIGCOMM 2019, Beijing, China, August 19-23, 2019, Jianping Wu and Wendy Hall (Eds.). ACM, 44–58. doi:10.1145/3341302.3342085 [51] Will Lin, Yizhou Shan, Ryan Kosta, Arvind Krishnamurthy, and Yiying Zhang. 2024. SuperNIC: An FPGA-Based, Cloud-Oriented SmartNIC. In Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays, FPGA 2024, Monterey, CA, USA, March 3-5, 2024, Zhiru Zhang and Andrew Putnam (Eds.). ACM, 130–141. doi:10.1145/3626202.3637564 [52] Linux RDMA. 2024. perftest – RDMA Performance Tests. https: //github.com/linux-rdma/perftest. Accessed: 04/15/2026. [53] Junyi Liu, Aleksandar Dragojević, Shane Fleming, Antonios Katsarakis, Dario Korolija, Igor Zablotchi, Ho-Cheung Ng, Anuj Kalia, and Miguel Castro. 2024. Honeycomb: Ordered Key-Value Store Acceleration on an FPGA-Based SmartNIC. IEEE Trans. Comput. 73, 3 (2024), 857–871. doi:10.1109/TC.2023.3345173 [54] Ming Liu, Tianyi Cui, Henry Schuh, Arvind Krishnamurthy, Simon Peter, and Karan Gupta. 2019. Offloading distributed applications onto
smartNICs using iPipe. In Proceedings of the ACM Special Interest Group on Data Communication, SIGCOMM 2019, Beijing, China, August 19-23, 2019, Jianping Wu and Wendy Hall (Eds.). ACM, 318–333. doi:10.1145/ 3341302.3342079 [55] Rui Ma, Evangelos Georganas, Alexander Heinecke, Sergey Gribok, Andrew Boutros, and Eriko Nurvitadhi. 2022. FPGA-Based AI Smart NICs for Scalable Distributed AI Training Systems. IEEE Computer Architecture Letters 21, 2 (2022), 49–52. doi:10.1109/LCA.2022.3189207 [56] MangoBoost. 2025. Mango BoostX™ Programmable DPUs. https: //cdn.sanity.io/files/hx87iaks/production/ce5454fc6af423cd241b5784 3750527b05d29811.pdf. Accessed on 04/15/2026. [57] YoungGyoun Moon, SeungEon Lee, Muhammad Asim Jamshed, and KyoungSoo Park. 2020. AccelTCP: Accelerating Network Applications with Stateful TCP Offloading. In 17th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2020, Santa Clara, CA, USA, February 25-27, 2020, Ranjita Bhagwan and George Porter (Eds.). USENIX Association, 77–92. https://www.usenix.org/confere nce/nsdi20/presentation/moon [58] Javier Moya, Matthias Gabathuler, Mario Ruiz, and Gustavo Alonso. 2023. fpgasystems/hacc: ETHZ-HACC. Zenodo. doi:10.5281/zenodo.8 340448 https://doi.org/10.5281/zenodo.8340448. [59] NVIDIA. [n. d.]. GPUDirect RDMA and GPUDirect Storage — NVIDIA GPU Operator. https://docs.nvidia.com/datacenter/cloud-native/gpuoperator/25.3.1/gpu-operator-rdma.html#gpudirect-rdma-andgpudirect-storage [60] NVIDIA. 2023. NVIDIA BlueField-3 DPU Data Sheet. https://www.nv idia.com/content/dam/en-zz/Solutions/Data-Center/documents/dat asheet-nvidia-bluefield-3-dpu.pdf. [61] NVIDIA Corporation. 2024. NVIDIA DOCA SDK. https://developer.nv idia.com/networking/doca Version 2.6.0, Accessed: 2026-03-24. [62] Oracle. 2025. Oracle Unveils Next-Generation Oracle Cloud Infrastructure Zettascale10 Cluster for AI. Oracle Corporation. https: //www.oracle.com/news/announcement/ai-world-oracle-unveilsnext-generation-oci-zettascale10-cluster-for-ai-2025-10-14/ Retrieved March 25, 2026. [63] Sourav Panda, Yixiao Feng, Sameer G Kulkarni, K. K. Ramakrishnan, Nick Duffield, and Laxmi N. Bhuyan. 2021. SmartWatch: accurate traffic analysis and flow-state tracking for intrusion prevention using SmartNICs. In Proceedings of the 17th International Conference on Emerging Networking EXperiments and Technologies (Virtual Event, Germany) (CoNEXT ’21). Association for Computing Machinery, New York, NY, USA, 60–75. doi:10.1145/3485983.3494861 [64] Charles Papon. 2016. SpinalHDL Documentation. https://spinalhdl.gi thub.io/SpinalDoc-RTD/master/SpinalHDL/Introduction/SpinalHD L.html. Accessed: 2025-04-15. [65] Suchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena, and Matthew D. Sinclair. 2023. Tale of Two Cs: Computation vs. Communication Scaling for Future Transformers on Future Hardware. In IEEE International Symposium on Workload Characterization, IISWC 2023, Ghent, Belgium, October 1-3, 2023. IEEE, 140–153. doi:10.1109/IISWC5 9245.2023.00026 [66] Salvatore Pontarelli, Roberto Bifulco, Marco Bonola, Carmelo Cascone, Marco Spaziani Brunella, Valerio Bruschi, Davide Sanvito, Giuseppe Siracusano, Antonio Capone, Michio Honda, and Felipe Huici. 2019. FlowBlaze: Stateful Packet Processing in Hardware. In 16th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2019, Boston, MA, February 26-28, 2019, Jay R. Lorch and Minlan Yu (Eds.). USENIX Association, 531–548. https://www.usenix.org/confe rence/nsdi19/presentation/pontarelli [67] Benjamin Ramhorst, Dario Korolija, Maximilian Jakob Heer, Jonas Dann, Luhao Liu, and Gustavo Alonso. 2025. Coyote v2: Raising the Level of Abstraction for Data Center FPGAs. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13-16, 2025, Youjip 14
[78] Marco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer, Lucian Petrica, Dario Korolija, Marco D. Santambrogio, Davide Conficconi, Gustavo Alonso, and Kenneth O’Brien. 2026. RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAs. In Proceedings of the Twenty-First European Conference on Computer Systems (Edinburgh, Scotland) (EuroSys ’26). Association for Computing Machinery, New York, NY, USA. [79] Zeke Wang, Hongjing Huang, Jie Zhang, Fei Wu, and Gustavo Alonso. 2022. FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUs. In Proceedings of the 2022 USENIX Annual Technical Conference, USENIX ATC 2022, Carlsbad, CA, USA, July 11-13, 2022, Jiri Schindler and Noa Zilberman (Eds.). USENIX Association, 967–986. https://www.usenix .org/conference/atc22/presentation/wang-zeke [80] Xingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen, and Haibo Chen. 2023. Characterizing Off-path SmartNIC for Accelerating Distributed Systems. In 17th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2023, Boston, MA, USA, July 10-12, 2023, Roxana Geambasu and Ed Nightingale (Eds.). USENIX Association, 987–1004. https://www.usenix.org/conference/osdi23/presentation/weismartnic [81] Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2022, Renton, WA, USA, April 4-6, 2022, Amar Phanishayee and Vyas Sekar (Eds.). USENIX Association, 945– 960. https://www.usenix.org/conference/nsdi22/presentation/weng [82] Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma, and Aleksandar Kuzmanovic. 2024. Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads. In 2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 767–784. https://www.usenix.org/conferenc e/atc24/presentation/xiao [83] Xilinx/AMD. 2024. Alveo Collective Communication Library (ACCL). https://github.com/Xilinx/ACCL. Accessed: 2024-05-22. [84] Jie Zhang, Hongjing Huang, Lingjun Zhu, Shu Ma, Dazhong Rong, Yijun Hou, Mo Sun, Chaojie Gu, Peng Cheng, Chao Shi, and Zeke Wang. 2023. SmartDS: Middle-Tier-centric SmartNIC Enabling Applicationaware Message Split for Disaggregated Block Storage. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA ’23). Association for Computing Machinery, New York, NY, USA, Article 42, 13 pages. doi:10.1145/3579371.3589077 [85] Yiran Zhang, Qingkai Meng, Chaolei Hu, and Fengyuan Ren. 2024. Revisiting Congestion Control for Lossless Ethernet. In 21st USENIX Symposium on Networked Systems Design and Implementation, NSDI 2024, Santa Clara, CA, April 15-17, 2024, Laurent Vanbever and Irene Zhang (Eds.). USENIX Association, 131–148. https://www.usenix.org /conference/nsdi24/presentation/zhang-yiran [86] Zhen Zhang, Chaokun Chang, Haibin Lin, Yida Wang, Raman Arora, and Xin Jin. 2020. Is Network the Bottleneck of Distributed Training?. In Proceedings of the 2020 Workshop on Network Meets AI & ML, NetAI@SIGCOMM, Virtual Event, USA, August 14, 2020, Behnaz Arzani and Xin Jin (Eds.). ACM, 8–13. doi:10.1145/3405671.3405810 [87] Zhipeng Zhao, Hugo Sadok, Nirav Atre, James C. Hoe, Vyas Sekar, and Justine Sherry. 2020. Achieving 100Gbps Intrusion Prevention on a Single Server. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 1083–1100. https: //www.usenix.org/conference/osdi20/presentation/zhao-zhipeng [88] Guanwen Zhong, Aditya Kolekar, Burin Amornpaisannon, Inho Choi, Haris Javaid, and Mario Baldi. 2023. A Primer on RecoNIC: RDMAenabled Compute Offloading on SmartNIC. CoRR abs/2312.06207 (2023). arXiv:2312.06207 doi:10.48550/ARXIV.2312.06207
Won, Youngjin Kwon, Ding Yuan, and Rebecca Isaacs (Eds.). ACM, 639–654. doi:10.1145/3731569.3764845 [68] Mario Ruiz, David Sidler, Gustavo Sutter, Gustavo Alonso, and Sergio López-Buedo. 2019. Limago: An FPGA-Based Open-Source 100 GbE TCP/IP Stack. In 29th International Conference on Field Programmable Logic and Applications, FPL 2019, Barcelona, Spain, September 8-12, 2019, Ioannis Sourdis, Christos-Savvas Bouganis, Carlos Álvarez, Leonel Antonio Toledo Díaz, Pedro Valero-Lara, and Xavier Martorell (Eds.). IEEE, 286–292. doi:10.1109/FPL.2019.00053 [69] Rob Rydberg, Madison N. Emas, John Demme, Ana Ibarra, Kara Kagi, Brandon Klouchek, Abhijeet Lawande, Todd Massengill, David J. Powers, and Andrew Putnam. 2026. Hyperscale FPGA Engineering Systems at Microsoft. In Proceedings of the 2026 ACM/SIGDA International Symposium on Field Programmable Gate Arrays, FPGA 2026, Seaside, CA, USA, February 22-24, 2026, Jing Li and Grace Zgheib (Eds.). ACM, 147–157. doi:10.1145/3748173.3779203 [70] Leah Shalev, Hani Ayoub, Nafea Bshara, and Erez Sabbag. 2020. A Cloud-Optimized Transport Protocol for Elastic and Scalable HPC. IEEE Micro 40, 6 (2020), 67–73. doi:10.1109/MM.2020.3016891 [71] Rajath Shashidhara, Tim Stamler, Antoine Kaufmann, and Simon Peter. 2022. FlexTOE: Flexible TCP Offload with Fine-Grained Parallelism. In 19th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2022, Renton, WA, USA, April 4-6, 2022, Amar Phanishayee and Vyas Sekar (Eds.). USENIX Association, 87–102. https: //www.usenix.org/conference/nsdi22/presentation/shashidhara [72] David Sidler, Zeke Wang, Monica Chiosa, Amit Kulkarni, and Gustavo Alonso. 2020. StRoM: smart remote memory. In EuroSys ’20: Fifteenth EuroSys Conference 2020, Heraklion, Greece, April 27-30, 2020, Angelos Bilas, Kostas Magoutis, Evangelos P. Markatos, Dejan Kostic, and Margo I. Seltzer (Eds.). ACM, 29:1–29:16. doi:10.1145/3342195.3387519 [73] Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan M. G. Wassel, Naveen Kr. Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar, Behnam Montazeri, Neelesh Bansod, Sarin Thomas, Inho Cho, Hyojeong Lee Seibert, Baijun Wu, Rui Yang, Yuliang Li, Kai Huang, Qianwen Yin, Abhishek Agarwal, Srinivas Vaduvatha, Weihuang Wang, Masoud Moshref, Tao Ji, David Wetherall, and Amin Vahdat. 2025. Falcon: A Reliable, Low Latency Hardware Transport. In Proceedings of the ACM SIGCOMM 2025 Conference, SIGCOMM 2025, São Francisco Convent, Coimbra, Portugal, September 8-11, 2025, Marília Curado, Christian Esteve Rothenberg, George Porter, and Srikanth Kandula (Eds.). ACM, 248–263. doi:10.1145/3718958.3754353 [74] Athinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai, Saksham Agarwal, Jacob Adelmann, David Ahern, Carlo Contavalli, Michael D. Goldflam, Vitaly Mayatskikh, Raghu Raja, Daniel Walton, Rachit Agarwal, Shrijeet Mukherjee, and Christos Kozyrakis. 2024. High-throughput and Flexible Host Networking for Accelerated Computing. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, Ada Gavrilovska and Douglas B. Terry (Eds.). USENIX Association, 405–423. https://www.usenix.org/conference/osdi24/presentation/sk iadopoulos [75] Shangyi Sun, Rui Zhang, Ming Yan, and Jie Wu. 2022. SKV: A SmartNIC-Offloaded Distributed Key-Value Store. In IEEE International Conference on Cluster Computing, CLUSTER 2022, Heidelberg, Germany, September 5-8, 2022. IEEE, 1–11. doi:10.1109/CLUSTER51413.2022.00 016 [76] Neil C. Thompson and Svenja Spanuth. 2021. The decline of computers as a general purpose technology. Commun. ACM 64, 3 (2021), 64–72. doi:10.1145/3430936 [77] Nathan Tibbetts, Sifat Ibtisum, and Satish Puri. 2026. A survey on heterogeneous computing using SmartNICs and emerging data processing units. Future Gener. Comput. Syst. 176 (2026), 108207. doi:10.1016/J.FUTURE.2025.108207
15
[89] Yibo Zhu, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, and Ming Zhang. 2015. Congestion Control for Large-Scale RDMA Deployments. In Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication, SIGCOMM 2015, London, United Kingdom, August 17-21, 2015, Steve Uhlig, Olaf Maennel, Brad Karp, and Jitendra Padhye (Eds.). ACM, 523–536. doi:10.1145/2785956.2787484 [90] Noa Zilberman, Yury Audzevich, G. Adam Covington, and Andrew W. Moore. 2014. NetFPGA SUME: Toward 100 Gbps as Research Commodity. IEEE Micro 34, 5 (2014), 32–41. doi:10.1109/MM.2014.61
16