A Protocol-Independent Transport Architecture Kimiya Mohammadtaheri University of Waterloo
Matthew Chen
David Gao
University of Waterloo
Eric Su
University of Waterloo
University of Waterloo
Saad Syed
Chris Neely
Mario Baldi
AMD
Nachiket Kapre
arXiv:2605.02210v1 [cs.NI] 4 May 2026
Pengyu Ji
University of Waterloo University of Waterloo
none
Mina Tahmasbi Arashloo
University of Waterloo
University of Waterloo
workloads such as Amazon’s SRD, Google’s Falcon, STRACK, and the Ultra Ethernet Consortium’s UET [15, 17, 24, 26]. Evolving the transport layer, however, is becoming increasingly difficult. To deliver high throughput (≥ 100𝐺𝑏𝑝𝑠) and low latency (tens of 𝜇𝑠) for modern data center workloads, while minimizing host CPU overhead, key transport functionality is increasingly implemented directly in network interface cards (NICs) [2, 15, 23, 24, 26]. Such hardware acceleration enables efficient data transfers at line rates of 100𝐺𝑏𝑝𝑠 and beyond. However, it also hardens the transport layer: existing hardware transports are either closed-box designs or expose programmable components that are limited in scope or expressiveness. As a result, modifying or deploying new transport functionality often requires modifying low-level hardware designs, making hardware transport far less adaptable than the rapidly evolving workloads it serves. As such, recent work has explored ways to increase the programmability of hardware transport. One category exposes programmability only in the control path while keeping the protocol logic implemented in the data path fixed (e.g., Falcon [26]). Specifically, the data path implements a fixed transport protocol that dictates how per-flow state evolves and how core transport operations such as packet generation, loss detection, recovery, and data reassembly are performed at line rate. The control path receives a predefined set of aggregated statistics from the data path and adjusts a fixed set of data-path parameters, such as pacing rates or timeout values. While the control algorithm can be modified, the protocol logic in the data-path remains fixed. That is, changing the core transport logic from one protocol to another, for example, evolving RoCE to Falcon, or Falcon to UET, would require modifying the low-level data path hardware design. Another category targets data-path programmability. Unlike the control path, the data path must operate under stringent performance constraints: since it executes stateful transport logic that produces and consumes packets, it must support line rates of hundreds of millions of packets per second. As such, designing a programmable transport datapath requires carefully balancing flexibility and performance. Existing systems navigate this tradeoff by incorporating built-in
Abstract The network transport layer is increasingly implemented in the NIC hardware to meet the performance demands of modern workloads, but this has made it difficult to evolve or deploy new transport protocols. Existing approaches either fix protocol logic in the data-path or build protocol-specific assumptions into the architecture that limit the range of protocols that can be supported on a single hardware substrate. We present PITA, a protocol-independent transport architecture that enables full data-path programmability while sustaining line-rate performance. PITA eliminates protocolspecific assumptions by structuring the data-path around a uniform abstraction over events, state, and instructions, and rethinks core components, including scheduling, packet generation, and data reassembly, to operate on this abstraction. We evaluate PITA along key dimensions reflecting the goals of its protocol-agnostic datapath design. Specifically, we show that PITA supports diverse protocol semantics by showing it can implement TCP and RoCEv2 on the same data path and preserve their distinct end-to-end behavior. Through targeted microbenchmarks and synthesis on Alveo U250 cards, we show that PITA’s redesigned components sustain high performance under demanding conditions, with modest hardware overhead and meeting timing at 250MHz.
1
Samuel Zhang
University of Waterloo
Introduction
The network transport layer sits on the critical path of every networked application. It receives data transfer requests from applications and determines how to reliably and efficiently deliver data over a shared and possibly unreliable network. Because it determines the performance achieved by the network and consequently perceived by applications, the transport layer is repeatedly extended or redesigned as applications, workloads, and network environments evolve. Beyond the many TCP variants and optimizations over the years [1, 12, 16, 27, 28], the past decade has seen a steady stream of new transport designs, including receiver-driven protocols for data centers [6, 10, 13, 20], RDMA over Converged Ethernet (RoCE) [11], and transports tailored to emerging 1
Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
assumptions about protocol behavior in their architecture. These assumptions constrain which parts of the protocol state are accessible at different decision points, the amount and type of computation that can be performed in response to events, and the kinds of events that can be processed. While effective for achieving high efficiency, this pushes the transport data path toward one end of the flexibility–performance tradeoff, constraining the range of transport protocols that can be supported on the same hardware substrate (§2). Protocol-Independent Transport Architecture. In this paper, we propose a hardware architecture for the network transport layer that enables full data-path programmability while sustaining line-rate performance. Our architecture is inspired by a recent high-level abstraction for network transport protocols [19], which represents the semantics of a transport protocol as mappings from user-defined events and per-flow state to an updated state and a sequence of protocolagnostic transport instructions. This abstraction suggests a more uniform structure across transport protocols and motivates revisiting the design of the transport data path to remove protocol-specific assumptions in how events, state, and protocol actions are represented and processed. Our Protocol-Independent Transport Architecture (PITA) consists of the common components of a transport datapath: event ingestion and scheduling, event processing pipelines, and modules for packet generation, data reassembly, and timer management. Each component is carefully designed to be either protocol-agnostic with minimal configuration or fully programmable, avoiding partially programmable modules with protocol-specific assumptions, and the interfaces between them follow a protocol-independent abstraction inspired by [19]. Specifically, event ingestion and scheduling, and the packet generation, reassembly, and timer modules are implemented through protocol-agnostic mechanisms that operate on generic events and instructions, and the event processing pipeline is fully programmable to express protocol-specific logic within the same abstraction. Realizing this design required revisiting each component of the transport data path, as removing protocol-specific assumptions introduces distinct challenges and design considerations across the system. For example, for event ingestion and scheduling, PITA needs to ingest one incoming event and dispatch one safe-for-processing event per cycle while ensuring state consistency under generic event streams, without relying on any protocol-specific structure or relationships between events or their processing. PITA addresses this through lightweight coordination between the scheduler and event processors to track event eligibility and enforce state consistency, while minimizing concurrent accesses to the event store buffers and metadata structures. Similarly, for packet generation, the absence of protocolspecific assumptions requires an instruction-driven design in which instructions fully specify how data is segmented and placed in packets. Sustaining a continuous packet output
stream then requires per-instruction data prefetching and efficient interleaving of instructions based on their parameters. Data reassembly is similarly instruction-driven, managing per-flow reassembly buffers under arbitrary segment placement and according to instruction parameters governing when reassembled data is made available to the application. Summary of contributions. We revisit the design of hardware transport datapaths to eliminate protocol-specific assumptions while preserving high performance. In doing so, PITA rethinks both the overall data-path architecture and the design of its individual components,such as event scheduling, packet generation, and reassembly, to operate over generic events, state, and instructions, and shows that a single hardware substrate can efficiently support diverse transport semantics. Together with prior work on controlpath programmability, this paves the way towards a fully programmable network transport layer in the NIC. Evaluation highlights. We evaluate PITA along three dimensions that reflect the key goals of its protocol-agnostic data-path design. First, we show that PITA supports diverse protocol semantics by programming it to implement TCP and RoCEv2, two protocols with radically different semantics, and show that PITA preserves their distinct end-to-end behavior under induced congestion. Second, through targeted microbenchmarks, we demonstrate that the redesigned scheduler, packet generator, and reassembly modules sustain high performance under demanding conditions. Finally, we evaluate PITA’s timing and resource utilization on AMD Alveo U250 FPGAs to show that it achieves these capabilities with modest hardware overhead and meets timing at 250MHz. Prototype will be open-sourced after publication.
2
Motivating Examples
Transport protocols manage reliable and efficient data transfer between endpoints over a shared, possibly unreliable network. Data is divided into segments, identified by their offset within a finite message or a bytestream, and transmitted in individual packets to the receiver. The receiver acknowledges received segments to help the sender track which data has arrived successfully. The sender and receiver cooperate to determine which segments and how many at a time should be (re)transmitted to ensure reliable and fast delivery without overwhelming the network and the receiver. Transport protocols have stateful and event-driven data paths. They keep per-“flow” state for a large number of flows, where the definition of a flow depends on the protocol, e.g., a connection in TCP, a queue-pair in RoCEv2, or a remote procedure call (RPC) in Homa. Protocol decisions are triggered by events such as application requests, packet arrivals, and timeouts, and involve updating and analyzing protocol state that summarizes the status of in-flight segments. To process events at high speed while maintaining state consistency, existing flexible hardware transport data paths 2
incorporate built-in assumptions about protocol behavior in their architecture. These assumptions constrain which parts of the protocol state are accessible at different protocol decision points in the data path, the amount and type of computation a protocol can perform when reacting to events, and the kinds of events a protocol can receive. The following examples illustrate several ways in which these assumptions appear in existing architectures. Example 1. In Tonic [2] and NanoTransport [3], the protocol-specific behavior in response to incoming control packets such as ACKs must, respectively, fit within a single pipeline stage (i.e., a single clock cycle, 4𝑛𝑠 at 100𝐺𝑏𝑝𝑠) or a feed-forward P4 pipeline whose stages cannot share state and support only a small number of semantically constrained read-modify-write operations. While this is sufficient for processing simple cumulative ACKs, it cannot support Selective ACKs (SACKs) which are increasingly incorporated into stream-based and message-based protocols to provide richer delivery information and efficiently recover from multiple close-together packet losses [18, 26]. Reacting to SACKs requires multiple dependent state reads, updates, and scans over bitmaps. Such operations exceed the stateful computation supported by Tonic and NanoTransport. In fact, the Tonic paper and code repository report implementing only a simplified variant that supports a single SACK range using a single find-first-set bitmap operation. Example 2. The data paths in Tonic and NanoTransport are designed around a fixed packet-generation strategy for data packets that constrains how protocol state can be accessed and exposed to the packet-generation logic. Specifically, both architectures maintain a bitmap in the perflow state that records only which segments are pending (re)transmission. A fixed-function module then uses this bitmap to asynchronously generate data segments, which is followed, in NanoTransport, by a P4 pipeline to manipulate packet headers. When reacting to incoming control packets such as ACKs, protocol logic can only convey information to the packet-generation path through this bitmap. The packet-generation logic and subsequent pipelines cannot access certain protocol state, such as receive-side information about which packets have been received so far. As such, these architectures are not expressive enough for certain common transport behaviors such as piggybacking control metadata about received packets onto outgoing data packets, like RPC completion signals in Homa or ACKs in TCP. Example 3. F4T [4] designs a hardware datapath specifically optimized for accelerating TCP. It observes that, for many TCP variants, incoming event metadata often affects per-flow state in simple ways: metadata values either replace existing state variables or update them using associative operations, such as addition, that can be implemented as a single-cycle read-modify-write. F4T leverages this observation by accumulating the “side effects” of TCP event metadata before invoking the main processing pipeline. This design
Figure 1. PITA’s architecture supports protocols with radically different semantics (§3) (green: fully programmable, purple: protocol-agnostic and reconfigurable) implicitly assumes that event effects on protocol state can be summarized using such simple associative updates. While this assumption holds for many TCP variants, it does not generalize to protocols in which events require more complex state interactions. For example, merging selective ACKs (SACKs), described above, may require multiple dependent state updates. Similarly, message-oriented protocols such as RoCEv2 or Homa treat each application request as an independent operation with its own parameters and completion state. Unlike TCP send requests, these events cannot be merged and must be tracked individually. Takeaways. Existing flexible hardware transport data paths incorporate built-in assumptions about protocol behavior in their architecture, and as such, restrict the range of transport protocols that can be implemented on the same hardware substrate.
3
PITA Overview
To provide full data-path programmability without embedding protocol-specific assumptions in the architecture, PITA follows the abstraction introduced in a recent domainspecific language (DSL) for the transport layer [19]. This abstraction represents the semantics of a transport protocol as mappings from an event and the per-flow state to an updated state and a sequence of transport instructions. Crucially, in this abstraction, events and per-flow state can include arbitrary, user-defined metadata. Moreover, the transport instructions that perform packet generation, data reassembly, and timer management are designed to be protocolindependent. They are parametrized to capture the axes along which protocols differ, e.g., header fields and segmentation rules for packet generation, allowing different transport protocols to be expressed using the same instruction set without embedding protocol-specific assumptions. Finally, the mappings from events and state to instructions are expressed 3
Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
as chains of simple C-like functions with bounded loops and no pointers. That is, they are only generic constraints that do not impose assumptions about protocol semantics. Protocol-independent architecture. By aligning PITA’s design with the above abstraction, the same hardware substrate will be able to support a wide range of transport protocols (e.g., TCP and RoCEv2) without embedding protocolspecific assumptions or logic in the datapath. Figure 1 provides a high-level view of PITA’s architecture. PITA’s modules operate on generic events and per-flow contexts, and a small set of pre-defined, configurable, and protocol-independent transport instructions. That is, the event-ingestion and scheduling pipeline accepts generic events generated by applications, packets, or timers, while the context table maintains userdefined per-flow state. The programmable event-processing pipelines in the protocol logic engine transform events and contexts into updated contexts and transport instructions, and the instruction-execution modules carry out packet generation, reassembly, and timer operations without making protocol-specific assumptions about how or when packets should be generated or reassembled in response to events. Programming PITA. To realize a particular transport protocol on PITA, the user configures a subset of its modules to define how events are parsed, what per-flow context is maintained, and how event-processing pipelines use the user-defined event metadata and context to generate transport instructions. We use TCP and RoCEv2 as representative examples here (and in our evaluation, §8.1) to demonstrate that PITA can support radically different transport semantics. PITA will naturally accommodate variations within protocol families as well, such as reprogramming features within TCP, RDMA-based, or RPC-based transport. To implement TCP, the user configures the event parsers in the event ingestion pipeline to extract the metadata required for TCP processing from incoming packets and socket send/receive requests. The context table is configured with the per-connection state required for TCP processing, such as an integer tracking the first sent but unacked data segment or a bitmap tracking received segments. The user also configures the timer module with TCP timers to handle lost data or control packets. The user then programs event-processing pipelines, one for each event. Each pipeline takes the incoming event and per-flow context as input and produces an updated context together with transport instructions for packet generation, data reassembly, or timer operations. For example, the user can program the TCP acknowledgment pipeline to compute how many bytes the sender may transmit based on window sizes and the new ACK information, and emit a packet-generation instruction specifying TCP header values, the starting data address, total bytes to send, the maximum segment size, and rules for updating sequence numbers during segmentation. The user can reconfigure the same architecture to implement a radically different protocol, such as RoCEv2. Unlike
Figure 2. PITA’s protocol-agnostic event scheduler (§4)(yellow: DP RAMs, blue: registers)
TCP, which manages reliable transfer of a byte stream using a sliding window, RoCE supports message-oriented RDMA operations between queue pairs and uses very different mechanisms for loss detection and recovery and congestion control. Nevertheless, the same event-processing infrastructure and instruction-execution pipelines support both protocols. Specifically, the user can reconfigure the event parsers to now extract metadata from RoCE packets and RDMA work queue elements (WQEs) instead and configure the context table with per–queue-pair state such as message sequence numbers and completion metadata rather than TCP bytestream state. Similarly, the event-processing pipelines can be reprogrammed to implement RoCE semantics for RDMA read, write, and other operations, emitting the same kinds of instructions but with different parameters, such as packet-generation instructions with appropriate packet and message sequence numbers and payloads, timer instructions to periodically trigger rate adjustments, and instructions to reassemble messages and generate WQE completion notifications. Realizing a single hardware substrate that efficiently supports such diverse protocols required revisiting the design of key data-path components, including event scheduling, packet generation, and reassembly. In §4 and §6, we describe how these components are designed to operate over generic events, state, and instructions. §7 discusses practical considerations, including enforcing atomic execution through backpressure, integrating PITA into a full transport stack with multiple data-paths and per-flow resource management, and opportunities for further optimization.
4
Event Ingestion and Scheduling
Transport protocols react to events originating from three sources: packets arriving from the network, application requests for data transfer, and timer expirations. Among these, packet arrivals are the most frequent and may occur at line rate, while application request rates vary across applications 4
depending on how network-intensive their execution is, and timer events are typically less frequent. To support a wide range of transport protocols without assuming specific event semantics or aggregation behavior, PITA’s event-ingestion and scheduling pipeline treats each event as a generic individual entity and is designed to sustain high throughput. In particular, the pipeline aims to ingest one event per cycle and dispatch one event per cycle to the protocol logic engine whenever at least one eligible event exists. An event is eligible for processing only if doing so preserves per-flow state consistency. Specifically, events belonging to the same flow must be processed in order, and a new event from a flow cannot enter the processing pipeline while a previous event from that flow is still being processed. To enforce these constraints while maintaining high throughput, PITA organizes incoming events into per-flow FIFO queues that are maintained in an event store and uses an event scheduler that tracks flows with eligible events and selects one each cycle (Figure 2). Implementing this design efficiently requires careful management of memory accesses to the event queues and the metadata structures in the event store and scheduler to minimize concurrent read and write accesses to each stateful module while sustaining one event insertion and one event dispatch per cycle. Efficient tracking of eligible events. Because events belonging to the same flow must be processed in order, PITA schedules events at the granularity of flows. A flow is considered eligible when (1) its event FIFO is non-empty and (2) no previous event from that flow is currently in the eventprocessing pipeline. Each cycle, the scheduler selects an eligible flow ID and requests the corresponding event from the event store, which dequeues and returns the event at the head of that flow’s event queue. Tracking eligibility efficiently requires determining when a flow no longer has an event in the pipeline and whether additional events remain queued for that flow. Since each flow can have at most one event in the pipeline at a time, one approach would be to associate each flow with a counter initialized to the pipeline depth and decremented every cycle until the event is guaranteed to have exited. The scheduler could then consult the queue occupancy metadata in the event store to determine whether the flow becomes eligible again. However, this requires maintaining unnecessary counter metadata and update logic as well as additional accesses to the event-store metadata structures. Instead, PITA uses a combination of three lightweight mechanisms: returning events to signal pipeline completion, piggybacking queue-state information on dequeued events, and maintaining small per-flow flags to handle concurrent arrivals. Specifically, each event “returns” to the scheduler after completing pipeline processing, implicitly indicating that the flow no longer has an event in flight. When an event returns, the scheduler must determine whether another event from that flow is waiting in the queue. The event
store already maintains per-flow queue occupancy metadata to detect full queues, which is accessed on both event insertion and dequeue. Rather than reading this metadata again, PITA piggybacks this information on the dequeued event by tagging it with a last-event bit indicating whether it was the final queued event for that flow. This information may become stale if new events for the same flow arrive while the event is being processed. As such, PITA maintains two additional per-flow flags that record whether new events have arrived since the flow last became ineligible. Using the last-event bit on the returned event and these flags, the scheduler determines whether the flow should be reinserted into the eligible-flow set. Managing event buffers. PITA’s event store maintains per-flow event FIFOs using three dual-ported RAMs. One RAM stores the event metadata in a pool of ring buffers (one per flow), while two additional RAMs maintain the head and tail pointers respectively, allowing event insertion and dequeue to proceed independently at line rate. Specifically, given a ring-buffer index, event insertion reads the current tail pointer, writes the event into the corresponding location in the buffer memory, and updates the tail pointer. Event dequeue follows a similar process using the head pointer. While queue occupancy could in principle be derived from the head and tail pointers, doing so would require additional pointer arithmetic and break the independence between accesses to the head and tail pointers. Instead, PITA maintains occupancy incrementally using per-flow counters updated on every insertion and dequeue. As a result, all updates remain lightweight increments and decrements, and accesses to the head and tail pointers remain independent. Protocol-specific configurations tailor the behavior of generic modules to the user’s target protocol. For the event store and the scheduler, which are designed to process generic events, the users only need to provide the maximum expected event width for the target protocol. To extract protocol-specific event metadata, PITA includes a programmable parser placed before the event-ingestion and scheduling pipelines. The parser follows an architecture similar to those used in programmable switches [5] and allows users to populate protocol-specific event fields from incoming packets and application requests.
5
The Protocol Logic Engine
When an event exits the scheduling pipeline described in §4, PITA retrieves the state associated with the event’s flow from a context table. The context table stores per-flow state in a dual-ported RAM. Similar to the event-ingestion and scheduling pipelines, the context table remains protocolagnostic and only needs to know the width of the per-flow state. PITA then delivers the event and the corresponding per-flow context to the Protocol Logic Engine (PLE), which applies protocol-specific logic to them. 5
Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
Conceptually, the PLE realizes the abstraction described in §3: it maps an input event and context to an updated context together with a sequence of transport instructions. The generated instructions are then executed by the instructionexecution modules (§6) to do packet generation, data reassembly, or timer operations. As this mapping differs widely across protocols, PITA exposes a fully programmable PLE to the users while keeping the surrounding datapath infrastructure protocol-agnostic and configurable through generic parameters. Moreover, PITA keeps PLE’s input and output interfaces generic and protocol-independent, so that user-defined protocol logic can seamlessly integrate with the protocolagnostic scheduler and instruction-execution modules. In our implementation, the protocol logic engine is programmed using High-Level Synthesis (HLS) [7]. Users write a C++ program with a particular structure, and the HLS toolchain generates a pipelined hardware implementation in a hardware description language like Verilog that can be plugged into the rest of the architecture. In the PLE program, the user defines a set of event processing functions, one for each event type, that describe how the event interacts with and updates the context of its flow and which instructions should be generated as a result. To interact properly with PITA’s protocol-agnostic modules, the event processing functions follow the interface shown below:
Figure 3. PITA’s protocol-agnostic packet generation (§6.1). (green: reprogrammable, yellow: DP RAMs, blue: registers) range of transport protocols without embedding protocolspecific assumptions, these modules operate on generic instruction formats and configurable parameters rather than protocol-specific logic and provide efficient, reusable implementations of common transport operations. 6.1
Transport protocols differ in how they construct packets. Some, such as TCP, assume that all data for a connection resides in a single contiguous bytestream and generate packets by segmenting the available data from that stream into packets carrying byte offsets. Others, such as RoCEv2, generate packets for a sequence of RDMA operations, where each operation may access different and unrelated memory locations; packet generation must therefore gather data from arbitrary addresses while maintaining a shared packet sequence space across the queue pair. Protocols such as Homa organize packet generation at the granularity of individual RPCs, with each RPC handled independently rather than as part of a continuous stream or operation sequence. Instruction-driven packet generation. In PITA, packet generation makes no assumptions about header formats, data layout in memory, or how packets are derived from larger data units. Instead, consistent with the abstraction described in §3, it is driven entirely by the packet generation instructions emitted by the protocol logic engine. Each instruction carries all the information required to generate one or more packets, specifically the location and length of the data to be transmitted, the header to attach, and the parameters needed to update the header across segments, and scheduling and pacing parameters such as per-flow rate or credit. Figure 3 shows a high-level overview of PITA’s packet generation module, which organizes execution around per-flow instruction queues, per-instruction data pre-fetch, and interleaved packet construction across flows. Here, the definition of a flow is configurable to match the protocol’s needs. For example, it may correspond to a TCP connection, a RoCEv2 queue pair, or an individual RPC message in Homa. For each flow, the packet generator maintains a queue of instructions and tracks the progress of the currently active instruction.
typedef hls::stream stream; void X_event_processor( stream<X>& event_in, stream<ctx>& ctx_in, // inputs stream<ctx>& ctx_out, // updated context // instruction streams stream<pktGen_instr> p[N], stream<timer_instr> t[K], stream<reassm_instr> r[M]);
All inputs and outputs are expressed using the hls::stream type, which the HLS toolchain synthesizes into AXI4-Stream interfaces. Event and context metadata are delivered to the protocol logic engine as generic bit sequences. Users can interpret these bit sequences as protocol-specific metadata by defining appropriate C++ structs. Similarly, instruction streams use protocol-agnostic metadata structures describing the parameters required for packet generation, reassembly, and timer operations (§6). After processing, the updated context is written back to the context table, and the event, together with the last-event flag, returns to the scheduler to signal that processing for that flow has completed.
6
Flexible Packet Generation
Protocol-Agnostic Instruction Execution
This section describes the design of the dedicated modules that execute the transport instructions for packet generation (§6.1), data reassembly (§6.2), and timer operations (§6.2) that are generated by the protocol logic engine. To support a wide 6
Moreover, for each instruction, the packet generator maintains a data buffer that is continuously replenished with the payload data from the instruction’s specified memory locations. When an instruction is selected, the packet generator incrementally constructs packets from it according to the instruction parameters. If the instruction requires multiple packets, it remains active across multiple iterations until all data has been transmitted. To sustain high throughput, the packet generator interleaves instruction execution across flows in a round-robin fashion. Continuous payload pre-fetch. Data that goes into packet payloads is often stored in external memory. PITA avoids stalling on external memory accesses by decoupling data retrieval from packet construction through continuous pre-fetching. For each instruction, the packet generator maintains a dedicated payload buffer that is populated by a data-fetch engine. When an instruction arrives, data fetching begins immediately to retrieve payload data from the specified memory locations and store it in the buffer. The fetch engine continues to replenish the buffer in the background as packets are generated. Specifically, further in the pipeline, packet construction incrementally consumes data from this buffer. If the buffer’s available data falls below a predefined threshold, the data fetch engine is notified for additional fetches, hiding memory access latency and allowing data retrieval and packet construction to proceed concurrently. Instruction arbitration. The packet generator organizes instructions in per-flow ring buffers implemented using dualported RAMs (similar to §4) and schedules packet transmission at flow granularity. For each flow, PITA maintains the state of the currently active instruction along metadata such as execution progress and optional pacing parameters, specified in the instruction, in registers. This information is used by a pacing module to schedule instructions for packet construction. To avoid head-of-line blocking, an instruction can be preempted after generating a configurable amount of data or packets to return from packet construction, update its execution state, and be rescheduled by the pacing module. Customizable segmentation and packet construction. The packet constructor receives the parameters of the selected instruction from the scheduler, consumes payload data from the per-instruction buffer, and generates packets according to the specified segmentation parameters. For each packet, it reads one segment of pre-fetched data (with segment size defined by the instruction), attaches the specified header, and transmits the resulting packet. If the instruction parameters describe more data than fits in a single packet, the constructor iteratively produces multiple packets, consuming buffered data and triggering additional data fetches as needed to maintain continuous execution. A key component of this process is a configurable headerupdate module that determines how header fields evolve across segments. Rather than assuming a fixed update rule, the module can be configured to apply protocol-specific rules
that take in the current header, bytes transmitted so far from this instruction, and instruction parameters, and produce the next header. For example, in TCP, this component can be configured to add the instruction segment size to the sequence number of the current header to derive the sequence number for the next packet. For RoCE, it can be configured to set opcodes depending on whether the packet is the first, the middle, or the last one in an operation. 6.2
Data Reassembly and Timers
Incoming data packets may arrive out of order or with gaps, and transport protocols differ in how they determine ordering and when data is complete and ready for application consumption. Consistent with the abstraction in §3, PITA does not make any assumptions about ordering or completion. Instead, the protocol logic engine (PLE) explicitly specifies these semantics through reassembly instructions. Instruction-driven reassembly. When data packets pass through the programmable parser (§4), their payloads are temporarily stored in a ring buffer in memory in arrival order. For each packet, the address of its corresponding payload in that memory is provided to the PLE as part of the event metadata. Based on protocol-specific logic, the PLE determines the correct placement of each segment and issues an add-data-seg instruction with the address of the segment in the temporary memory and the correct offset as parameters. The reassembly module then fetches the payload from the temporary memory and inserts it at the specified offset in a per-flow reassembly buffer. The reassembly module also supports a flush-and-notify instruction, which makes a contiguous portion of the buffer available to the application. In both cases, the module performs only the operations specified in the instruction, without tracking segment state or inferring ordering, leaving all protocol semantics to the PLE. Handling segment alignment. Executing the add-dataseg instruction involves fetching the segment from temporary payload memory and inserting it into the reassembly buffer in a dual-ported RAM at the specified byte offset. Reassembly buffers are organized in memory-addressable fixed-size chunks. Because PITA does not assume any alignment guarantees from instructions, the reassembly module must support arbitrary byte offsets. To do that, it aligns incoming data with chunk boundaries using a shift pipeline and performing read-modify-write for the first and last partial chunks when needed. This process is largely pipelined, where inserting a segment spanning 𝑁 chunks takes 𝑁 + 1 cycles. The extra cycle arises from misalignment and boundary reads, needed to support arbitrary offsets. Larger chunk sizes require supporting a wider range of shifts; we use 64B chunks, matching the minimum packet size. For example, consider inserting a 256B segment at offset 71 into a buffer with 64B chunks. The shift amount is given by the offset modulo the chunk size, i.e., 71 mod 64 = 7. The segment spans four chunks, which are passed through a shift 7
Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
pipeline that applies shifts based on the set bits in the bit representation of the shift amount (here, shifting by 1, 2, and 4 bytes). Starting from the second chunk, the pipeline concatenates the previous and current chunks, shifts them together, and selects the first 64 bytes of the result. This ensures that any leftover data from shifting the previous chunk is correctly accounted for. For the first and potentially last partial chunks, existing buffer contents must be preserved. In this example, the first 7 bytes of the first chunk and the last 57 bytes of the final chunk are read from memory, merged with the shifted data, and written back. Flush and notify. Once the protocol logic engine determines that a contiguous region of data of size 𝑋 is ready, it issues an instruction to transfer 𝑋 more bytes from the reassembly buffer to an application-provided memory address. To support this, the reassembly module maintains a per-flow read pointer into the buffer and, upon receiving a flush-and-notify instruction, outputs the required number of chunks starting from this pointer and advances it so that it always points to the first chunk with data not yet exposed to the application. This allows the reassembly module to provide the required data for transfer to the application solely based on instruction parameters and without interpreting protocol-specific completion conditions. Timer instructions. Protocols differ in the number of per-flow timers and when they start, restart, and stop them. PITA’s design allows users to specify the number of per-flow timers and manage them through instructions issued by the protocol logic engine. These timers rarely need to operate at a granularity finer than a few microseconds, and their management follows a design similar to prior work [2].
7
withholds events from affected flows, maintaining a queue of otherwise eligible flows. These flows are re-enabled once there is less downstream queue buildup, ensuring events are issued only when their instructions can be fully executed. Managing per-flow resources. PITA maintains per-flow resources across components, including scheduler event buffers, context table state, and instruction/data buffers in execution modules. Flow identifiers in events and instructions are mapped to indices identifying the corresponding per-flow resources in each component, which must be dynamically managed in long-running systems. The current prototype statically provisions these mappings to focus on demonstrating the programmability of the datapath. However, PITA is compatible with standard dynamic allocation techniques. For queue-based structures such as event and instruction buffers, a freelist-based allocator can be used to assign buffers to active flows, with buffers entering a draining state upon deallocation and returned to the freelist once empty. Per-flow state can be swapped to and from off-chip memory (e.g., DRAM) to support larger working sets. Incorporating these mechanisms into PITA requires modifying some the scheduling and resource management logic. However, since those rely on protocol-agnostic information such as flow identifiers and generic queue operations, they can be incorporated into PITA without introducing protocolspecific assumptions into the data path. Further optimizations. While PITA operates on finegrained events and instructions, it can support domainspecific optimizations such as event and packet coalescing. Such coalescing is not universally applicable across protocols (§2), and is therefore not embedded in PITA’s base architecture. However, consistent with [19] and analogous to segmentation rules, the scheduler and packet generator can expose interfaces for specifying protocol-specific coalescing policies. Moreover, §3–§6 describe a single transport data path. To scale further, multiple PITA instances can be used as parallel datapaths with a load balancer assigning flows to instances [4]. Integrating PITA into such systems is an interesting avenue for future work.
Practical Considerations
We have described how PITA achieves protocol-agnostic data-plane programmability by combining generic event scheduling (§4), a programmable protocol logic engine (§5), and protocol-independent instruction execution modules (§6). Together, these components provide a flexible substrate for implementing a wide range of transport protocols. In this section, we discuss additional practical considerations. Ensuring atomic event processing. Once an event exits the scheduler, it must be processed atomically. That is, if it updates flow state in the PLE, its generated instructions must not be dropped downstream. This could happen despite PITA’s components being pipelined because one event may generate an instruction whose execution spans multiple cycles. This multi-cycle execution is inherent to transport data paths: packet generation or data insertion into reassembly buffers involve data movement and are constrained by I/O bandwidth (e.g., transmitting a 1500B packet takes 24 cycles for 64B bus width). To ensure atomic processing, instruction execution modules provide backpressure signals when their buffers exceed a threshold. The scheduler then temporarily
8
Evaluation
We evaluate PITA along key dimensions that reflect the goals of its protocol-agnostic data-path design: (1) Support for diverse protocol semantics using TCP and RoCEv2 as representative examples (§8.1), (2) The ability of the redesigned core components to sustain efficient, line-rate operation (§8.2), and (3) Timing and resource overheads in hardware (§8.3). Implementation and evaluation setup. PITA’s prototype (to be open-sourced) is implemented in 5367 lines of System Verilog code. Users program the PLE using HLS, and the resulting event pipelines are plugged in with the other core components. All results are obtained using the AMD 8
User-defined events HLS LoC Pipeline Depth (Max) Pipeline Depth (Avg) FF Usage LUT Usage
TCP (w/ AIMD)
RoCEv2 (w/ DCQCN)
4 612 12 4.75 13384 (< 1%) 7346 (< 1%)
17 1434 6 2.35 132731 (3.8%) 51972 (3%)
Table 1. TCP and RoCEv2 implementation in PITA’s protocol logic engine. Each pipeline maps an event and input flow context to a sequence of instructions and the updated context (§5). Despite radically different semantics and structure, both are realized within the same protocol-agnostic datapath.
Figure 4. Validating faithful realization of TCP and RoCEv2 in PITA by comparing their end-to-end behavior for a keyvalue store application under induced congestion (§8.1). and 24KB (∼200 and 400 minimum-sized packets, respectively). For RoCEv2, the ECN marking threshold for DCQCN is set to 3KB (∼ 50 minimum-sized packets). To induce congestion, we temporarily reduce the queue drain rate to half the request generation rate, causing queue buildup before restoring the original rate. This high-load, small-message setting combined with induced congestion and the possibility of bufferbloat creates conditions under which the two protocols manifest different latency behavior, allowing us to validate that PITA faithfully captures their semantics. Figure 4 shows the resulting request–response latency for TCP and RoCEv2 under the two buffer sizes. As expected, RoCEv2 achieves lower average and tail latency (p90) under congestion, with the gap increasing at the larger queue size. This reflects the different transport mechanisms in the two protocols and shows that PITA preserves protocol-specific behaviors on the same protocol-agnostic hardware substrate.
Vivado toolchain [8], including its cycle-accurate simulator, which models pipeline and memory behavior at cycle granularity, at a target frequency of 250MHz. 8.1
Supporting Diverse Transport Semantics
To demonstrate PITA ’s ability to support diverse protocol semantics on a single hardware substrate, we program it to implement TCP and RoCEv2 (w/ DCQCN). These two protocols have radically different semantics and together cover a broad range of common mechanisms used in transport protocols that shape data-path behavior, including stream- vs. messageoriented semantics, different segmentation and packetization strategies, different kinds of acknowledgements, different pacing mechanisms (window vs. rate) (see §3 for details). Programming PITA involves configuring the main datapath components and implementing the event processing logic using HLS, as described in §3 and §5. While all components are configured per protocol, we focus on the HLS implementation within the PLE, which is responsible for mapping userdefined events and flow context to transport instructions and updated state according to protocol logic. Table 1 summarizes the results. The two protocols differ significantly in the number of events, complexity of the event-processing logic, and pipeline depths, and how they parameterize transport instructions, reflecting their distinct semantics. For example, RoCEv2 requires a larger set of events and more complex processing due to its operation-based model and richer packet semantics, while TCP relies on a smaller set of events but deeper processing pipelines. Despite these differences, both protocols are implemented within the same protocolagnostic datapath and incur modest hardware overhead. To validate faithful realization of the main mechanisms in these protocols, we evaluate their end-to-end behavior using a simple key-value client-server application. The client issues back-to-back requests with 4B keys at 8M Req/s, and the server responds with 64B values. The round-trip time is set to 6.4𝜇s, and we configure two queue sizes of 12KB
8.2
Sustaining Line-Rate in Core Components
Event ingestion and scheduling. To sustain line rate, PITA’s protocol-agnostic event scheduler needs to ingest one event and dispatch one event per cycle to the PLE if there is at least one eligible event whose processing will not violate state consistency (§4). Two factors primarily influence the scheduler’s event throughput: the PLE pipeline depth and event arrival patterns. PLE pipeline depth impacts how long a flow remains ineligible after issuing an event, and event arrival patterns, particularly the number of active flows and burstiness, affect the availability of eligible events. These factors are especially important during cold start, when the event store is initially empty and the scheduler must build up enough parallelism across flows to sustain line rate. Impact of PLE depth. To evaluate the impact of PLE depth, we generate events for 1024 flows, where each flow produces bursts of 10 consecutive events, and and measure the scheduler’s output throughput (averaged over 50 cycles) as we vary the PLE pipeline depth. Figure 5a shows how throughput approaches line rate from a cold start. For realistic PLE depths of 3 and 10 (TCP and RoCEv2 have maximum 9
Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
(a) Varying PLE pipeline depth
(b) Varying per-flow event burst size
(c) Impact of competing with other flows
Figure 5. PITA’s protocol-agnostic scheduler sustains line-rate throughput under realistic operating conditions (§8.2). While deep PLE pipelines (a) and extreme burstiness (b) can delay convergence from a cold start and increase intra-flow latency, they do not limit steady-state throughput, and the additional latency due to cross-flow contention (c) remains modest.
(a) B2B instructions, mixed packet size
(b) B2B instructions, fixed packet size
(c) B2B instructions, fixed segment size
Figure 6. PITA’s protocol-agnostic and instruction-driven packet generation and data reassembly sustains line rate under demanding conditions: back-to-back single-packet/segment instructions of various sizes. §8.2 discusses results and edge cases. pipeline depths of 12 and 7; Table 1), flows become eligible for rescheduling quickly, allowing the scheduler to rapidly approach one event per cycle. In contrast, with an extreme depth of 100, flows remain ineligible for longer, delaying the buildup of concurrency and increasing the time required to reach line-rate throughput from a cold start. Impact of event burst size. Figure 5b shows how the scheduler’s throughput approaches line rate from cold start for different per-flow burst sizes, 1024 flows and PLE depth of 3. Burstiness limits the number of flows with eligible events, reducing cross-flow parallelism, especially during cold starts. For a burst size of one (no burst), the scheduler reaches line rate immediately, because most flows do not have outstanding events in the pipeline (given the number of active flows relative to the PLE depth), making incoming events immediately eligible for dispatch. As burst size increases, throughput ramps up more slowly. This is most evident for a large burst size of 100: early on, only a few flows have events in the scheduler (one in the first 100 cycles, two in the first 200 cycles, and so on). For more realistic burst sizes (e.g., 10), the scheduler reaches line rate within ∼100 cycles. Scheduling latency. We use 1024 flows and PLE depth of 3 to measure how quickly events leave the scheduler after arrival. As burst size increases, overall scheduling latency increases, as events within a flow must wait for prior events from the same flow to complete processing. To separate this
intra-flow serialization latency from cross-flow contention, we define load-induced latency as the additional delay experienced by an event relative to the “ideal” scenario in which no other flows are present in the scheduler. Figure 5c shows the distribution of load-induced latency. We observe that this latency remains relatively small (∼10 cycles on average) and is largely insensitive to burst size. The lower latency observed for burst size one is due to minimal contention, described above, where events are immediately eligible upon arrival. Reducing intra-flow latency via techniques such as protocolagnostic but configurable event coalescing is an avenue for future work (§7) and fits well within PITA’s architecture. Event scheduler takeaways. PITA’s protocol-agnostic scheduler sustains line-rate throughput under realistic operating conditions. While extreme burstiness and deep PLE pipelines can delay convergence from a cold start and increase intra-flow latency, they do not limit steady-state throughput, and the additional latency due to cross-flow contention remains modest. Packet generation. PITA ’s packet generation is entirely instruction-driven: each instruction emitted by the PLE specifies the header, data, and segmentation parameters required to generate one or more packets. To evaluate whether the packet generator sustains line-rate operation, we disable rate limiting and generate instructions that each request a single 10
packet with a random size between 64B and 1500B. In practice, instructions typically generate multiple packets, and flows are paced by congestion control (§8.1), making this a stress test that exercises the packet generator under more demanding conditions than typical workloads. Figure 6a shows that at a 250MHz clock frequency, the packet generator comfortably sustains 100Gbps line rate under this workload. Figure 6b shows the packet generation rate under a sequence of single-packet instructions with fixed packet sizes 64B to 1500B. PITA sustains 100Gbps line rate for packet sizes of 128B and above. For minimum-sized packets (64B), the current prototype does not fully sustain line rate under back-to-back instructions. This is because transmission of a 64B packet completes in one cycle, requiring the packet constructor to wait for the next instruction for updated header and segmentation parameters. For larger packets, this latency is hidden by packet transmission time. This gap does not reflect a fundamental limitation and can be addressed using standard techniques such as instruction lookahead or prefetching. Moreover, this case arises only when flows issue back-to-back instructions, each corresponding to a single minimum-sized packet. Such patterns are uncommon in well-formed transport protocols, which, instead of issuing back-to-back individual small packets, batch data into multiple full-sized packets. Data reaasembly. As data packets arrive, the PLE issues add-data-seg instructions that place segments at the correct offsets in per-flow reassembly buffers. To sustain line-rate operation, the reassembly module must therefore consume segments at line rate. We evaluate this by generating a sequence of add-data-seg instructions with fixed segment sizes ranging from 64B to 1500B. Figure 6c shows that the design sustains 100Gbps line rate for segment sizes of 256B and above. For smaller segments (128B and below), the current prototype does not fully sustain line rate under back-to-back instructions. This is due to an additional cycle required when segment offsets are not aligned with the 64B chunk boundaries of the reassembly buffer (§6.2), which introduces occasional read-modify-write overhead. In practice, workloads rarely consist solely of back-to-back small segments, and mixed traffic allows this overhead to be amortized, enabling the reassembly module to sustain line-rate operation. 8.3
Flows 128 256 512 1024
LUT
E = 16 FF BRAM LUT
(1.7M) (3.5M) (2688)
E = 32 FF BRAM
(1.7M) (3.5M) (2688)
18K 24K 36K 57K
29K 35K 46K 71K
18K 24K 38K 60K
29K 35K 47K 72K
235 312 496 835
235 343 527 898
Table 2. Resource utilization on an AMD Alveo U250 card across flow counts and per-flow event buffer depth (𝐸). All meet timing at 250MHz.
and keep the data-path busy as long as buffers are non-empty. As a result, the main scalability knobs we vary are the number of flows and the depth of per-flow event buffers. Table 2 summarizes the resource utilization. All configurations meet timing at 250MHz. Despite supporting protocolagnostic execution, PITA incurs modest overhead, using <3.5% of LUTs and <2% of FFs. BRAM usage ranges from 8% for 128 flows to ∼32% for 1024. The higher usage of BRAM compared to other resources reflects the stateful nature of transport processing and is also observed by prior work [2, 4]. In practice, a single data path instance is expected to handle a smaller number of active flows (closer to 128 than 1024), with techniques such as load balancing across multiple data paths used to support larger workloads (§7, [4]).
9
Related Work
We discussed flexible hardware transport datapaths such as NanoTransport, Tonic, and F4T in §2. Here, we situate PITA within the broader categories of related approaches. Protocol-specific hardware transport. Several works implement transport protocols in FPGAs or ASICs [15, 23, 24, 26]. These designs are highly optimized for a specific protocol and adapting them to a different protocol typically requires substantial hardware modifications and deep implementation knowledge. In contrast, PITA provides a protocolagnostic datapath, allowing users to implement different protocols without modifying the underlying hardware. High-level synthesis for transport. EasyNet implements a full TCP stack in HLS [14]. While HLS offers a higher-level programming interface than RTL, EasyNet’s design remains tailored to TCP. Meeting the stringent performance requirements of transport datapaths requires writing HLS code with hardware constraints in mind (e.g., pipelining and memory layout), making it significantly different from conventional C++ code. As a result, extending these designs to support other protocols remains challenging and requires substantial re-engineering. SoC-based transport offloads. Some works offload parts of the transport stack to embedded processors in SoC-based
Timing and Hardware Resource Utilization
We synthesize PITA on an AMD Alveo U250 FPGA to evaluate its hardware cost and timing characteristics for 128-1024 flows and per-flow event buffer size of 16 and 32 (TCP in PLE, configuration details in Table 3 in §A). The dominant contributor to resource consumption is BRAM allocated for per-flow event, instruction buffers, and context. Instruction buffers can remain relatively shallow and are managed using backpressure (§7). This is because each packet-generation instruction typically produces multiple large packets, allowing packet transmission to overlap with instruction execution 11
Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
SmartNICs [21, 25], with the remaining functionality running on the host CPU. These designs offer a more softwarelike development model than FPGA or ASIC approaches. However, they typically cannot achieve the same line-rate performance as FPGA- or ASIC-based designs [4, 9]. Software transport programmability. Prior work has proposed transport-layer abstractions, primarily in software. CCP [22] focuses on congestion control, while MTP [19] provides a general abstraction for transport protocols. PITA builds on these abstractions and shows how they can inform the design of a fully programmable hardware datapath.
10
[8] Vivado Developers. [n. d.]. AMD Vivado™ Design Suite. https://www.amd.com/en/products/software/adaptive-socs-andfpgas/vivado.html. ([n. d.]). Accessed: January 2025. [9] Daniel Firestone, Andrew Putnam, Sambhrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian Caulfield, Eric Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitu Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, and Albert Greenberg. 2018. Azure Accelerated Networking: SmartNICs in the Public Cloud. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18). USENIX Association, Renton, WA, 51–66. https://www.usenix.org/conference/nsdi18/presentation/firestone [10] Peter X Gao, Akshay Narayan, Gautam Kumar, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker. 2015. pHost: Distributed near-optimal datacenter transport over commodity network fabric. In Proceedings of the 11th ACM Conference on Emerging Networking Experiments and Technologies. 1–12. [11] Chuanxiong Guo, Haitao Wu, Zhong Deng, Gaurav Soni, Jianxi Ye, Jitu Padhye, and Marina Lipshteyn. 2016. RDMA over commodity ethernet at scale. In Proceedings of the 2016 ACM SIGCOMM Conference. 202–215. [12] Sangtae Ha, Injong Rhee, and Lisong Xu. 2008. CUBIC: a new TCPfriendly high-speed TCP variant. ACM SIGOPS operating systems review 42, 5 (2008), 64–74. [13] Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew W Moore, Gianni Antichi, and Marcin Wójcik. 2017. Rearchitecting datacenter networks and stacks for low latency and high performance. In Proceedings of the Conference of the ACM Special Interest Group on Data Communication. 29–42. [14] Zhenhao He, Dario Korolija, and Gustavo Alonso. 2021. EasyNet: 100 Gbps Network for HLS. In 2021 31st International Conference on Field-Programmable Logic and Applications (FPL). 197–203. https://doi. org/10.1109/FPL53798.2021.00040 [15] Intersect360 Research. 2025. UEC 1.0: New High-Performance Standard for Scaling HPC-AI. White Paper. Ultra Ethernet Consortium. https://ultraethernet.org/wp-content/uploads/sites/20/2025/ 06/UEC1.0Whitepaper.pdf [16] Cheng Jin, David X Wei, and Steven H Low. 2004. FAST TCP: motivation, architecture, algorithms, performance. In IEEE INFOCOM 2004, Vol. 4. IEEE, 2490–2501. [17] Yanfang Le, Rong Pan, Peter Newman, Jeremias Blendin, Abdul Kabbani, Vipin Jain, Raghava Sivaramu, and Francis Matus. 2024. Strack: A reliable multipath transport for ai/ml clusters. arXiv preprint arXiv:2407.15266 (2024). [18] Radhika Mittal, Alexander Shpiner, Aurojit Panda, Eitan Zahavi, Arvind Krishnamurthy, Sylvia Ratnasamy, and Scott Shenker. 2018. Revisiting network support for RDMA. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication (SIGCOMM ’18). Association for Computing Machinery, New York, NY, USA, 313–326. https://doi.org/10.1145/3230543.3230557 [19] Pedro Mizuno, Kimiya Mohammadtaheri, Linfan Qian, Joshua Johnson, Danny Akbarzadeh, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo. 2026. A Target-Agnostic Protocol-Independent Interface for the Transport Layer. (2026). arXiv:cs.NI/2509.21550 https://arxiv.org/abs/2509.21550 [20] Behnam Montazeri, Yilong Li, Mohammad Alizadeh, and John Ousterhout. 2018. Homa: A receiver-driven low-latency transport protocol using network priorities. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication. 221–235. [21] YoungGyoun Moon, SeungEon Lee, Muhammad Asim Jamshed, and KyoungSoo Park. 2020. AccelTCP: Accelerating Network Applications
Conclusion
We present PITA, a protocol-independent transport architecture that removes protocol-specific assumptions from the data path while sustaining high performance. By structuring the data path around a uniform abstraction over events, state, and instructions, PITA demonstrates that flexibility and efficiency need not be at odds in hardware transport design. We believe this approach opens a path toward fully programmable transport layers in NICs, enabling rapid evolution of transport protocols to meet emerging workloads.
References [1] Mohammad Alizadeh, Albert Greenberg, David A Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan. 2010. Data center tcp (dctcp). In Proceedings of the ACM SIGCOMM 2010 Conference. 63–74. [2] Mina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford, David Walker, and David Wentzlaff. 2020. Enabling Programmable Transport Protocols in High-Speed NICs. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, Santa Clara, CA, 93–109. https://www. usenix.org/conference/nsdi20/presentation/arashloo [3] Serhat Arslan, Stephen Ibanez, Alex Mallery, Changhoon Kim, and Nick McKeown. 2021. NanoTransport: A Low-Latency, Programmable Transport Layer for NICs. In Proceedings of the ACM SIGCOMM Symposium on SDN Research (SOSR) (SOSR ’21). Association for Computing Machinery, New York, NY, USA, 13–26. https://doi.org/10.1145/ 3482898.3483365 [4] Junehyuk Boo, Yujin Chung, Eunjin Baek, Seongmin Na, Changsu Kim, and Jangwoo Kim. 2023. F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework. In Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA ’23). Association for Computing Machinery, New York, NY, USA, Article 55, 13 pages. https://doi.org/10.1145/3579371.3589090 [5] Pat Bosshart, Glen Gibb, Hun-Seok Kim, George Varghese, Nick McKeown, Martin Izzard, Fernando Mujica, and Mark Horowitz. 2013. Forwarding metamorphosis: fast programmable match-action processing in hardware for SDN. In Proceedings of the ACM SIGCOMM 2013 Conference on SIGCOMM (SIGCOMM ’13). Association for Computing Machinery, New York, NY, USA, 99–110. https://doi.org/10.1145/ 2486001.2486011 [6] Qizhe Cai, Mina Tahmasbi Arashloo, and Rachit Agarwal. 2022. dcPIM: Near-optimal proactive datacenter transport. In Proceedings of the ACM SIGCOMM 2022 Conference. 53–65. [7] Vitis Developers. [n. d.]. AMD Vitis HLS. https://www.amd.com/en/products/software/adaptive-socs-andfpgas/vitis/vitis-hls.html. ([n. d.]). Accessed: January 2025. 12
with Stateful TCP Offloading. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, Santa Clara, CA, 77–92. https://www.usenix.org/conference/nsdi20/ presentation/moon [22] Akshay Narayan, Frank Cangialosi, Deepti Raghavan, Prateesh Goyal, Srinivas Narayana, Radhika Mittal, Mohammad Alizadeh, and Hari Balakrishnan. 2018. Restructuring endpoint congestion control. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication (SIGCOMM ’18). Association for Computing Machinery, New York, NY, USA, 30–43. https://doi.org/10.1145/3230543. 3230553 [23] Mario Ruiz, David Sidler, Gustavo Sutter, Gustavo Alonso, and Sergio López-Buedo. 2019. Limago: An FPGA-Based Open-Source 100 GbE TCP/IP Stack. In 2019 29th International Conference on Field Programmable Logic and Applications (FPL). 286–292. https://doi.org/10. 1109/FPL.2019.00053 [24] Leah Shalev, Hani Ayoub, Nafea Bshara, and Erez Sabbag. 2020. A cloud-optimized transport protocol for elastic and scalable hpc. IEEE micro 40, 6 (2020), 67–73. [25] Rajath Shashidhara, Tim Stamler, Antoine Kaufmann, and Simon Peter. 2022. FlexTOE: Flexible TCP Offload with Fine-Grained Parallelism. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). USENIX Association, Renton, WA, 87–102. https: //www.usenix.org/conference/nsdi22/presentation/shashidhara [26] Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan MG Wassel, Naveen Kr Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar, Behnam Montazeri, Neelesh Bansod, et al. 2025. Falcon: A reliable, low latency hardware transport. In Proceedings of the ACM SIGCOMM 2025 Conference. 248–263. [27] Kun Tan, Jingmin Song, Qian Zhang, and Murad Sridharan. 2006. A compound TCP approach for high-speed and long distance networks.
In Proceedings-IEEE INFOCOM. [28] Balajee Vamanan, Jahangir Hasan, and TN Vijaykumar. 2012. Deadlineaware datacenter tcp (d2tcp). ACM SIGCOMM Computer Communication Review 42, 4 (2012), 115–126.
A
PITA Parameters
Table 3 shows the relevant PITA parameters. Module
Parameter
Value
Global
flow count event width event type count context width serialized data width serialized packet width
ED 64 b (TCP) 4 (TCP) 938 b (TCP) 64B 64B
Scheduler
per-flow event buffer depth eligible flows queue depth
ED = flow count
Pkt Gen
header width per-flow instr. queue depth per-flow pre-fetch buffer len.
168 b 8 64 × 64B
Reassembly per-buffer reassembly buffer len 256 × 64B
Table 3. Parameter values for results in §8.3. ED stands for experiment-dependent
13