Who Needs DRAM? We Have Fiber Hannah Atmer
Thiemo Voigt
Uppsala University Uppsala, Sweden
Uppsala University Uppsala, Sweden
Yuan Yao
Stefanos Kaxiras
Uppsala University Uppsala, Sweden
Uppsala University Uppsala, Sweden
arXiv:2607.08407v1 [cs.AR] 9 Jul 2026
Abstract
performed by every node that serves the same model. Replication leads to excessive energy consumption. Insight 2: Fiber is Memory. The sheer amount of fiber in a hyperscale datacenter (10,000 to 100,000 km of fiber strand) holds an immense number of bits at any time. A simple loop of fiber that spans the datacenter has a capacity of multiple TB of data that circulate past every compute node in the loop at 2/3 the speed of light (𝑐). Effectively, we can turn fiber into a Delay-line Memory, one of the first types of memory used in electronic computers [34], albeit with a tremendous speed, capacity, and length, compared to the Mercury delay-line memories of the 1940’s and early 1950’s. Instead of thousands of nodes consuming local electrical energy to fetch identical weights from HBM or DDR, a single centralized optical transmitter can stream the model parameters once into the shared fiber network. Inference nodes passively “tap” the data circulating continuously in the fiber, extracting weights on the fly and eliminating redundant weight storage and fetch energy across the cluster. Our proposal is supported by the growing adoption of co-packaged optics (CPO) which places silicon photonics engines directly onto the processor substrate to bypass powerhungry electrical transceivers [12]. In co-package optics, the silicon switch chip and the optical engines (silicon photonics chips) are placed on the same package substrate. The electrical signal only has to travel millimeters instead of centimeters, which avoids the need for a digital signal processing unit which consumes most of the power of the optical to electrical conversion [13] [27]. This capability obviates the need to buffering in RAM and allows the possibility of feeding data directly into compute units. Existing state-of-the-art datacenter networking hardware also seeks to minimize data transfer costs but is still based on intermediary buffering. For example, NVIDIA’s NVNetIO SmartNIC receives the incoming network packets and streams the raw data directly into the GPU’s VRAM using GPUDirect RDMA [23] and some FPGA SmartNICs briefly buffer data in pipeline registers, small FIFOs, or FPGA block RAM [31]. In our approach, however, we strive to obviate, not only the need for storage for immutable data, but also
The rising pressure on DRAM availability and contract pricing reflects generative AI’s massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expansion, which now consumes a significant portion of global DRAM output. In this work, we propose a new architecture: Fiber Memory, which reimagines the role of optical fiber in a hyperscale data center, deploying it as an active, recirculating delay-line memory for immutable data, such as large language model (LLM) weights. We present a data-parallel optical broadcast delayline memory architecture that accounts for fiber’s physical realities. By incorporating space-division multiplexed multi-core fibers (MCFs), passive optical tap-and-amplify interfaces, co-packaged optics (CPO), and regional all-optical regeneration, our case study evaluation demonstrates that Fiber Memory can eliminate redundant weight storage across 10,000 AI accelerators and reduce weight-delivery energy by over 70% compared to traditional HBM3e configurations.
1
Introduction and Motivation
Memory is a bottleneck in modern computing clusters running large language models (LLMs). Traditional hardware platforms depend on stacking high-bandwidth memory (HBM) or double-data-rate (DDR) DRAM directly adjacent to processing units to feed billions of model parameters into arithmetic pipelines. This paradigm has driven a massive surge in DRAM demand, resulting in supply constraints, high costs, and thermal and power limits within hyperscale data centers [26]. In this work, we propose a new paradigm where the optical fiber network is used as memory to avoid unnecessary data replication. Insight 1: Weight replication in DRAM memory across a datacenter is immensely inefficient. Consider that in a hyperscale datacenter, the same model parameters (e.g., Attention and MLP weights) are replicated across all the nodes that serve the same LLM. Not only model parameters are replicated extensively, but far worse, the accesses (requests-responses) to such replicated data are identically This work was supported by the Swedish Foundation for Strategic Research (SSF) grant FUS21-0067. 1
Hannah Atmer, Thiemo Voigt, Yuan Yao, and Stefanos Kaxiras
for any buffering on their way to the computing units where they will be used. In this work, we: (1) Establish the feasibility of using data center fiber as a high-speed delay-line memory for AI accelerators. (2) Sketch an optical network architecture utilizing multi-core fibers, passive splitters, and optical amplifiers to distribute weight streams without electronic conversions. (3) Detail data representation, alignment, and interleaving strategies to feed processing systolic arrays directly from fiber. (4) Provide a quantitative evaluation of the performance, latency, and energy consumption of Fiber Memory compared to standard HBM-based systems.
Weights Server
Splitter/ amplifier/repeater
Accelerator Pod 1
...
Accelerator Pod N
Figure 1: Ring and Pod Architecture. The weights server injects model parameters into a wavelengthmultiplexed 19-core MCF ring. Regional splitters broadcast the signal to independent pods, minimizing cumulative losses and local transceiver overhead.
2 Fiber as Memory 2.1 Bit Capacity of Fiber and the Bandwidth-Delay Product The capacity of an optical fiber to store data “in flight” is governed by the bandwidth-delay product (BDP). The propagation speed of light in a standard silica fiber core (𝑣) is given by:
99:1 Optical Splitter
99% 1%
1x100 Splitter
𝑐 3 × 108 m/s ≈ = 2 × 108 m/s = 200 km/ms (1) 𝑛 1.5 where 𝑛 ≈ 1.5 is the refractive index of the fiber core [8]. The total propagation delay (𝜏) for an optical fiber strand of length 𝐿 is: 𝐿 𝜏= (2) 𝑣 If a transmitter modulates the light at an aggregate data rate of 𝐵, the total volume of data 𝑀 stored inside the fiber strand at any single instant is: 𝐿 𝑀 =𝐵 ×𝜏 =𝐵 × (3) 𝑣 Aggregating the route lengths of a hyperscale data center’s fiber optic cables yields an aggregate fiber strand length (𝐿) ranging from 10, 000 to 100, 000 km [7] [36]. Using ultra-dense WDM, a single commercial fiber strand can achieve an aggregate bandwidth (𝐵) of approximately 100 Tb/s (12.5 TB/s). Substituting these metrics into the BDP equation reveals the enormous latent storage capability of the network. For a 100,000 km aggregate fiber length, the storage capacity is: 𝑣=
PDFA Optical Amplifiers
100-Accelerator Pod
99:1 Optical Splitter
99% 1%
1x100 Splitter
PDFA Optical Amplifiers
Optional 2R O-E-O Repeater
100-Accelerator Pod
Figure 2: Asymmetric Tap-and-Amplify Receiver Schematic. The passive 1 : 99 tap extracts 1% of the signal power for local execution while letting 99% pass through. Regional PDFAs restore signal amplitude at the Pod-level, minimizing active component counts.
While this example demonstrates the baseline potential of a single fiber strand, our complete Fiber Memory architecture aggregates bundles of multi-core cables to achieve a vastly higher total bandwidth and capacity. We utilize space-division multiplexed 19-core Multi-Core Fibers (MCFs). Packing 19 independent cores within a single physical glass cladding sharing a single protective jacket allows us to compress the spatial footprint of the delay line by over 90% [16], making the physical installation of the 1, 000 km loop highly manageable within standard datacenter cable trays. To exploit this in-flight storage while maintaining physical realizability, we organize the optical network into a ring and pod architecture (Fig. 1). A central weights server stores the weights in non-volatile memory and writes them onto the fiber ring at startup. Regional passive optical splitters then broadcast the optical signals to distinct local “Pods” of 10 inference chassis. Inside each Pod, the signal runs along a localized distribution bus.
100, 000 km = 6.25 TB (4) 200, 000 km/s This capacity is more than sufficient to store multiple massive LLMs (such as a 1-trillion parameter model quantized to INT8 or FP16) entirely in flight. 𝑀 = 12.5 TB/s ×
2
Who Needs DRAM? We Have Fiber
2.2
Receiver Hardware
Optical Engine PIC
Traditional network transceivers are point-to-point: they receive an optical signal, convert it to the electrical domain (O-E), process or route it, and re-modulate it back to light (E-O). This O-E-O cycle is highly energy-intensive and introduces hundreds of nanoseconds of digital processing latency. Instead, our architecture uses passive, highly asymmetric "Tap-and-Amplify" interfaces. As shown in Fig. 2, each chassis splits the incoming local distribution bus using a 1 : 99 passive optical splitter. • Branch A (The Tap): Diverges a tiny fraction of the optical power (1%) directly to the co-packaged optical receivers on the local AI accelerators. Because the distance from the splitter to the silicon photodetector is mere millimeters, this tap is practically instantaneous and introduces no electronic buffering. • Branch B (The Ring): Directs the remaining optical power (99%) back into the local bus to propagate to the next downstream chassis. Using a highly asymmetric 1 : 99 tap ensures that the through-path insertion loss per chassis is extremely small: Losstap = −10 log10 (0.99) ≈ 0.043 dB
Systolic Array
Connector
TSV-interposer Fiber Source
ASIC Substrate Host PCB
Figure 3: Co-Packaged Optics (CPO) accelerator integration. The Photonic Integrated Circuits (PICs) are positioned on the same interposer as the compute silicon. Passive micro-ring resonators (MRRs) demultiplex the optical wavelengths from 32 MCF cores and feed raw weight parameters directly to the systolic array registers.
chassis receives the full 14-cable MCF bundle (carrying all 256 active cores), allowing the chassis to ingest the entire 128 GB model stream simultaneously. This guarantees that each chassis is 100% independent and can complete end-toend inference passes without transferring activations over a backend network. However, terminating 256 physical fibers into a single silicon chip presents severe manufacturing and thermal challenges. To address this, we physically distribute the 256 optical cores internally among the 8 accelerators on the chassis board. Each individual accelerator package only physically taps a sub-group of 32 cores. Inside the local PIC of each accelerator (Fig. 3), an array of silicon micro-ring resonators (MRRs) demultiplexes only 8 wavelengths per core. This restricts the total on-chip receiver interface to just 256 MRRs (32 cores × 8 wavelengths), which is highly manufacturable, thermally stable, and commercially viable using current silicon photonics packaging techniques [27].
(5)
To amplify these multi-wavelength WDM streams in the O-band (1310 nm), we employ Praseodymium-Doped Fiber Amplifiers (PDFAs). PDFAs utilize a fluoride glass host doped with 𝑃𝑟 3+ ions to provide broad, stable gain across the 1280– 1330 nm window [20]. Importantly, PDFAs exhibit a long upper-state lifetime (≈ 110 μs), rendering them immune to the inter-channel cross-gain modulation (XGM) and fast gain saturation that plagues Semiconductor Optical Amplifiers (SOAs) in multi-wavelength setups [24]. This allows us to deploy only two PDFAs per Pod (one booster PDFA at the head of the Pod, and one inline preamplifier PDFA) per cable to offset all split, tap, and fiber losses. The PDFA amplifies the optical carrier waves directly in the optical domain, avoiding O-E-O conversion. The latency of a PDFA is determined solely by the time of flight through the doped fluoride fiber segment (typically < 15 m), which is less than 75 ns [19]. All-optical 2R regeneration (Re-amplification and Re-shaping) is deployed regionally to suppress noise accumulation without electrical conversion [17].
2.3
FAU
EIC
3 LLM Inference at Scale 3.1 LLM Weight vs. Activation Footprints LLM execution is highly asymmetric with respect to memory access patterns. During the autoregressive decoding phase of inference, the accelerator generates tokens sequentially, one by one. For each generated token, the processor must read the complete set of model weights (𝑊 ) from memory, while the activation data (𝐴) consists predominantly of the Key-Value (KV) cache for preceding tokens. The volume of weights vastly exceeds the size of the activations for small to medium batch sizes. In typical workloads, the model weights account for 90% to 99% of the total memory bandwidth consumption [3]. For instance, to generate a single token in a 70-billion parameter model (FP16, 140 GB size), a GPU must load the entire 140 GB of parameters into its registers, whereas the activation and KV cache transfers
Co-Packaged Optics and Pipeline-Parallel Integration
Data-Parallel Chassis Architecture. We define our foundational network “Node” not as a single monolithic chip, but as an 8-accelerator baseboard chassis (analogous to standard dense computing platforms like NVIDIA HGX). The entire 3
Hannah Atmer, Thiemo Voigt, Yuan Yao, and Stefanos Kaxiras Case A: Tight Back-to-Back Packing (No Slack) Layer 𝑁
Layer 𝑁 + 1
Case B: Empty Timing Slack Spaces
Figure 4: Structure of a streamed weight packet. The packet begins with a lock-synchronizing preamble, followed by a FEC-protected header containing the layer ID and scaling factors and then the payload containing raw weight elements unrolled to map to the accelerator’s spatial arrays, and a cyclic redundancy check.
Layer 𝑁
entire loop cycle (𝜏) for the weights to reappear. To build robustness and scheduling flexibility, the weights server can implement interleaved replication and padding. As shown in Fig. 5, we can use slack spaces (empty optical carrier windows or high-frequency idle patterns) between layer packets to give the node a safety margin for finalizing its activation and KV cache memory swaps before the next layer payload begins. Additionally, we can replicate and interleave weights at multiple points within the physical ring, reducing the worst-case wait time for a node seeking to begin a new inference cycle. However, both slack spaces and interleaving shrink the effective storage capacity of the fiber network.
Data Alignment and Systolic Array Interfacing
Because the weights are received continuously from the fiber, the compute units must execute matrix-vector multiplications in lockstep with the incoming light pulses. Standard processors request data using memory addresses, but Fiber Memory operates on a push-based, deterministic streaming model: the accelerator simply waits for the required layer parameters to flow through the tap. To structure this stream, weights are packaged into Streamed Weight Packets (SWPs), depicted in Fig. 4. Each SWP contains the following parts: Preamble: A highly distinct optical pulse pattern that allows the receiver’s clock and data recovery (CDR) circuits to lock onto the incoming data phase. Frame Header: The header contains information necessary for nodes to recover from operational jitter and stalls, make use of replication, and protects the quantization scaling factors. Payload: The raw weight matrix elements (𝑊𝑖 ), formatted specifically to match the processing layout of the systolic array. CRC: A trailing Cyclic Redundancy Check field used for final packet boundary verification. Since the weight matrix is read directly from the tap into the execution registers without local buffering, the weights server unrolls the matrix layout beforehand, thereby eliminating the need for complex address translation, rowdecoding, or layout transformations on the accelerator.
3.3
Layer 𝑁 + 1
Figure 5: Interleaving and padding comparison. In Case A, consecutive layer parameters are tightly packed; a slight computational delay causes a node to miss Layer 𝑁 + 1. In Case B, the inclusion of empty or redundant slack spaces allows nodes with operational jitter to safely synchronize and decode the incoming stream.
total less than 1.4 GB. Because LLM inference is memorybandwidth bound, performance is limited by how quickly the weights can be transferred from HBM. Under our paradigm, we keep the activations and KV cache in a small, local DRAM or high-density SRAM on the accelerator, while streaming 100% of the massive weight footprint from the fiber.
3.2
Timing Slack
4
Quantitative Llama-3-70B Case Study
To evaluate our architecture under physically consistent and realistic design constraints, we analyze a cluster deployment running a 70-Billion Parameter Dense Model (Llama-3-70B) quantized to INT8 (70 GB footprint). 4.0.1 Model and Cluster Parameters. We configure the system with the following metrics: Model Size (𝑊 ): 1 model = 70 GB (INT8), with an aggregate fiber capacity scaled to 128 GB to accommodate timing slack spaces and redundant layer replicas. The up to 45% slack space in this design accommodates the increase in attention latency associated with longer context lengths. Cluster Size: 1,250 independent 8Accelerator Chassis (totaling 10,000 accelerators). These are organized into 125 regional Pods (10 chassis per Pod). Space-Division Multiplexed Spooling: To store 128 GB (1.024 Terabits) with a physically consistent delay, we utilize 14 parallel multi-core links of 19-core Multi-Core Fiber (MCF), where each of the 14 links is composed of 20 cascaded stages of standard commercial 50 km spools. Across the 14 cables, this provides 266 total optical cores, leaving 10 cores unallocated for hot-swappable redundancy and keeping 256 active cores.
Replication, Interleaving, and Slack
A major challenge in a delay-line system is coordinating the processing rate of individual nodes with the constant speed of the optical stream. If a node falls behind due to a high batch size or prolonged KV cache lookup, it might miss the beginning of the next weight packet, forcing it to wait an 4
Who Needs DRAM? We Have Fiber
Direct-Detection Spectral Plan: To eliminate the need for power-hungry coherent DSPs [4], we exploit the natural zero-dispersion window of silica fiber in the O-band (𝜆0 ≈ 1312 nm). To prevent Inter-Symbol Interference (ISI) across the 1, 000 km run, we replace wide-band coarse division multiplexing with Dense Wavelength Division Multiplexing (DWDM) using tight 100 GHz (≈ 0.6 nm) channel spacing centered directly around 𝜆0 . Each of the 256 active cores carries 8 DWDM channels running 50 Gbaud PAM4 (100 Gb/s per wavelength), yielding a single-core bandwidth of 800 Gb/s (100 GB/s). Aggregate Link Bandwidth (𝐵): Across the 256 active cores, the aggregate bandwidth is 25.6 TB/s (204.8 Tb/s).
This 1024 kW calculation represents a theoretical maximum assuming continuous, peak-bandwidth utilization. Although practical inference workloads experience computational micro-stalls that slightly reduce dynamic fetch rates, we deliberately omit the substantial static leakage and refresh power overhead inherent to massive HBM3e clusters, keeping this baseline conservatively balanced. Furthermore, while Fiber Memory completely eliminates the energy cost of weight delivery, nodes will still expend a small fraction of local energy accessing on-chip SRAM or high-density DRAM for the remaining 1% to 10% of memory bandwidth required by activations and the KV cache. 4.0.5 Fiber Memory Energy and Power Projections. Central Transmitter & Lasers: We modulate 256 active cores × 8 channels = 2,048 total lasers. At 100 mW optical power per laser and an O-band Distributed Feedback (DFB) laser wallplug efficiency of 5% [11], the central laser source draws:
4.0.2 Physical Ring Dimensioning & Delay. The propagation delay (𝜏) of light traveling through the 1, 000 km spooled fiber loop is: 𝜏=
1, 000 km = 5 ms 200, 000 km/s
(6)
𝑃lasers_central = 2, 048 × 2 W ≈ 4.1 kW
Central Delay Line Amplification: The 1, 000 km loop requires inline amplification between each 50 km spool to prevent total signal extinction. With 14 cables and 20 stages, we deploy 280 inline PDFAs. At 25 W each:
At an aggregate bandwidth of 25.6 TB/s, the maximum volume of data 𝑀 stored in flight across the bundle is: 𝑀 = 25.6 TB/s × 0.005 s = 128 GB
(10)
(7)
This is sufficient to hold our 70 GB Llama-3-70B INT8 model, leaving 58 GB for timing slack and interleaved packet replicas.
𝑃loop_amps = 280 × 25 W = 7.0 kW
(11)
O-band PDFAs & All-Optical 2R Regenerators: With 125 Pods routing the full 14-cable bundle, we deploy 2 PDFAs per cable (one booster and one inline) and 1 regional alloptical 2R regenerator per cable in every Pod. PDFAs provide optical Re-amplification (1R) and require 25 W of electrical power per amplifier [29]. We employ all-optical Re-shaping (2R) using a non-linear Semiconductor Optical Amplifier (SOA) configured for cross-phase modulation, drawing 4 W for thermal bias [28]:
4.0.3 Inference Scheduling & Receiver Path. By utilizing a pure data-parallel chassis model, the full 25.6 TB/s model bandwidth is routed into every chassis. Internally, the 256 cores are partitioned among the 8 localized processing engines. Therefore, the weight-delivery bandwidth processed by each single localized accelerator chip is exactly: Bandwidthaccelerator = 32 cores × 8 wavelengths × 100 Gb/s
𝑃optical_network = 125 Pods× (28×25 W+14×4 W) = 94.5 kW (12) Local IM-DD PIC Receivers, Equalizers & FEC: By utilizing direct-detection PAM4 instead of coherent technology, we bypass power-hungry ADCs and massive coherent DSPs. The short-reach receiver electronics operate at an aggregate 0.7 pJ/bit [2]:
= 25.6 Tb/s = 3.2 TB/s (8) This matches the typical memory bandwidth of a targeted, high-end AI processing engine (e.g., NVIDIA H100 with 3.35 TB/s HBM3 [22]) and is well within the capabilities of standard silicon photonics. 4.0.4 Baseline Energy and Power Projections. We compare our direct-detection model against a standard HBM3e baseline. We deliver a weight throughput of 25.6 Tb/s per node comparable to an AI accelerator package equipped with two 1.6 TB/s HBM3e [18]. Using a realistic energy metric of 4.0 pJ/bit for HBM3e local memory fetches [15], the weightdelivery power for 10,000 traditional HBM3e cluster nodes is:
𝑃 receivers = 10, 000 × 25.6 Tb/s × 0.7 pJ/bit = 179.2 kW (13) Summing these terms yields the total Fiber Memory power consumption: 𝑃fiber_total = 4.1 kW+7.0 kW+94.5 kW+179.2 kW = 284.8 kW (14) Comparing the baseline and fiber models, Fiber Memory achieves a 72.1% reduction in total weight-delivery power (284.8 kW vs. 1, 024 kW) across the cluster, while completely eliminating the static leakage power, cooling overhead, and
𝑃HBM = 10, 000 × 25.6 Tb/s × 4.0 pJ/bit = 1024 kW (9) 5
Hannah Atmer, Thiemo Voigt, Yuan Yao, and Stefanos Kaxiras
high capital expense of storing 700 Terabytes of replicated static model weights in localized HBM3e stacks.
optical probe beam [32], resetting the OSNR before the signal is distributed to the 100 accelerators within the pod.
5
Chromatic Dispersion Management. To eliminate excessive chromatic dispersion, our architecture localizes operations within the O-band’s zero-dispersion regime. We deploy an 8channel Dense WDM grid with 100 GHz spacing, confining the entire transmission within a narrow 700 GHz (∼ 4.2 nm) spectral window centered exactly at 𝜆0 = 1312 nm. By restricting the total spectral footprint (Δ𝜆) and centering it precisely at the fiber’s zero-dispersion wavelength, we ensure the physical dispersion coefficient remains bounded at |𝐷 | ≤ 0.1 ps/(nm · km) across all active channels, minimizing pulse broadening without the need for active dispersion compensation [33].
Challenges and Physical Constraints
Deploying a multi-kilometer recirculating fiber memory requires careful management of the physical limitations of fiber optics. Fiber Attenuation and Splitter Losses. Standard single-mode fiber cores inside our Multi-Core Fiber operating in the Oband (1310 nm) exhibit a physical loss of ≈ 0.32 dB/km [25]. As detailed in Section 2.2, a 1 : 99 local tap introduces a tiny through-path insertion loss of 0.043 dB. For 10 chassis in a Pod, the cumulative tap loss is only ≈ .43 dB. A PDFA provides +20 dB to +30 dB of gain [21]. A PDFA that connects consecutive 50 km sections of fiber is located at the input of each Pod compensates for the distribution split and line losses, and a PDFA placed halfway through the Pod distribution bus offsets tap and local propagation losses.
Forward Error Correction and LLM Error Tolerance. To maximize the distance between expensive optical regenerators, we utilize high-throughput Forward Error Correction integrated directly into the accelerator’s PIC receiver. Low-overhead Reed-Solomon or Low-Density Parity-Check codes can correct raw input Bit Error Rates as high as 10−3 down to a clean 10−12 operating standard, with minimal latency overhead (< 100 ns [6]). Furthermore, LLMs exhibit resilience to minor noise. Quantized neural network weights are highly error-tolerant: random single-bit flips in the weight matrices have negligible impact on the model’s outputs, provided the scaling factors and layer normalizations remain uncorrupted [9]. By applying strong FEC protection specifically to the SWP headers and scaling factors, we can tolerate minor, uncorrected bit flips in the raw weight payloads, relaxing the OSNR and physical transceiver requirements of the network.
Amplification and O-Band Amplifier Noise Accumulation. Optical amplifiers introduce Amplified Spontaneous Emission (ASE) noise [10]. Each amplification step degrades the Optical Signal-to-Noise Ratio (OSNR) by adding random phase and amplitude fluctuations. The accumulation of ASE noise limits how far the signal can propagate before becoming unreadable. The OSNR after 𝑁 cascaded amplification stages can be approximated by standard analytical models [1]): OSNR𝑁 ≈
𝑃 in 𝑁 × (𝐺 − 1) × 𝑛 sp × ℎ𝜈 × Δ𝜈
where 𝑃in is the input signal power, 𝐺 is the amplifier gain, 𝑛 sp is the spontaneous emission factor of the Praseodymiumdoped fluoride core, ℎ𝜈 is the photon energy at 1310 nm, and Δ𝜈 is the optical bandwidth. To maintain an OSNR above the threshold required for low bit-error-rate detection—typically >15 dB for 50 Gbaud PAM4 formats to keep the pre-FEC BER below acceptable limits [35] — we must limit the consecutive analog amplification steps. The primary ASE noise contributors are the inline PDFAs connecting the twenty 50 km trunk segments (each requiring an amplifier gain 𝐺 ≈ 16 dB to offset the 0.32 dB/km O-band attenuation) and the distribution amplifiers serving the 10-chassis pods. Evaluating the OSNR approximation with these physical gain parameters reveals that the signal approaches the 15 dB threshold after circulating through the cascaded trunk segments and the localized 1 : 99 chassis taps. We route the stream through a regional all-optical 2R regenerator (Re-amplification and Re-shaping) at the pod level to prevent signal degradation. This device utilizes cross-phase modulation inside a non-linear SOA or Highly Non-Linear Fiber to transfer data cleanly onto a fresh
Thermal Management, Laser Reliability, and External Sources. Silicon photonics chips are highly sensitive to thermal variations. Micro-ring resonators (MRRs) rely on submicron dimensions to achieve wavelength resonance; a minor temperature shift drifts the refractive index, causing the MRR to miss its target wavelength channel [14]. Additionally, laser diodes are the most vulnerable component in any optical system. High-power laser diodes operated on hot processor interposers degrade rapidly, presenting a severe reliability risk [5]. To address these challenges, we utilize External Laser Sources. The high-power lasers are housed in separate, cool, hot-swappable chassis modules on the front panel of the server racks, far from the hot accelerator chips. The light is guided into the CPO package via passive optical fibers. Active thermal tuning is relegated to micro-thermoelectric heaters integrated into the MRRs, which consume only micro-watts per channel to stabilize the optical alignment [30]. 6
Who Needs DRAM? We Have Fiber
6
Conclusion and Future Outlook
[7] Corning Incorporated. 2025. Extreme-Density Ribbons and Multi-Core Fiber Engineering. Corning Optical Product Catalogs. [8] Fosco Connect. 2026. Optical Fiber Tutorial - Optic Fiber - Communication Fiber. https://www.fiberoptics4sale.com/blogs/archive-posts/ 95146054-optical-fiber-tutorial-optic-fiber-communication-fiber. [Online; accessed 29-June-2026]. [9] Zhen Gao, Qiang Liu, Jie Deng, et al. 2026. Enabling the Use of Approximate Memories in Quantized Large Language Models (LLMs): Dealing with Errors on the Scaling Factors. TechRxiv. Published February 27, 2026. [10] C. Randy Giles and Emmanuel Desurvire. 1991. Modeling ErbiumDoped Fiber Amplifiers. Journal of Lightwave Technology 9, 2 (1991), 271–283. doi:10.1109/50.65886 [11] Ibrahim Guler, Chen Sun, and Rajeev J. Ram. 2023. Laser Source Constraints and Wall-Plug Efficiency Optimization for High-Density Silicon Photonics Interconnects. Journal of Lightwave Technology 41, 8 (2023), 2411–2420. doi:10.1109/JLT.2023.3235612 [12] Tingbo He. 2026. A Time Scaling Theory for Multi-Layer Electronic Systems. ChinaXiv. https://chinaxiv.org/abs/202605.00224 [13] IEEE Electronics Packaging Society. 2026. Chapter 9: Photonics. Heterogeneous Integration Roadmap (HIR), 2026 Edition. IEEE EPS. https://eps.ieee.org/wp-content/uploads/2026/05/HIR_ 9_Photonics_rev0.9.pdf [14] Siegfried Janz, Sergey Dedyulin, D. X. Xu, Martin Vachon, Shurui Wang, Ross Cheriton, and John Weber. 2023. Measurement Accuracy in Silicon Photonic Ring Resonator Thermometers: Identifying and Mitigating Intrinsic Impairments. arXiv (2023). doi:10.48550/arxiv. 2306.15643 [15] Kyomin Lee, Junyoung Kim, Taeyoung Oh, Kyoung-Hoi Sihn, et al. 2024. A 1.2-TB/s HBM3E DRAM with Advanced Thermal Compression Non-Conductive Film Packaging and Optimized Core Power Delivery. In 2024 IEEE International Solid-State Circuits Conference (ISSCC). 312– 314. doi:10.1109/ISSCC49657.2024.10454389 [16] Takashi Matsui, Taiji Sakamoto, Fumihiro Hanzawa, Shigeru Tomita, Kyozo Tsujikawa, and Kazuhide Nakajima. 2017. Low-Loss and LowDMD 6-Mode 19-Core Fiber With Cladding Diameter of Less Than 250 𝜇m. Journal of Lightwave Technology 35, 3 (2017), 443–449. doi:10. 1109/JLT.2016.2608920 [17] Antonio Mecozzi, Mark Shtaif, and Cristian Antonelli. 2018. AllOptical 2R and 3R Regeneration in Regional Optical Fiber Communication Systems. Journal of Lightwave Technology 36, 7 (2018), 1411–1420. doi:10.1109/JLT.2017.2789104 [18] Micron Technology, Inc. 2024. Micron HBM3E Product Brief: Introducing Memory Built for AI Innovation. Product Brief. Micron Technology, Inc. https://assets.micron.com/adobe/assets/urn:aaid:aem:b710d8f2-7f6644c1-a234-456e2b986347/renditions/original/as/hbm3e-productbrief.pdf [19] Takashi Mori, Kazumasa Takada, and Takemi Hasegawa. 2022. UltraLow Latency O-Band Praseodymium-Doped Fiber Amplifiers for Data Center Interconnects. Journal of Lightwave Technology 40, 8 (2022), 2445–2452. doi:10.1109/JLT.2022.3141108 [20] Takashi Mori, Kazumasa Bro Takada, Hanawa Nishi, and Takemi Hasegawa. 2021. O-Band Praseodymium-Doped Fluoride Fiber Amplifiers for Multi-Wavelength High-Speed Data Center Interconnects. Journal of Lightwave Technology 39, 12 (2021), 3932–3940. doi:10.1109/JLT.2021.3061223 [21] M. Nishi, K. Takada, T. Mori, and T. Hasegawa. 2022. High-Gain OBand Praseodymium-Doped Fluoride Fiber Amplifiers for High-Speed Data Center Interconnects. Journal of Lightwave Technology 40, 14 (2022), 4689–4696. doi:10.1109/JLT.2022.3164401 [22] NVIDIA. 2026. NVIDIA H100 Tensor Core GPU. https://www.nvidia. com/en-us/data-center/h100/ Accessed: 2026-07-08.
We have presented Fiber Memory, a new architecture that repurposes datacenter fiber networks as active, looping delayline memory systems. By streaming static, highly replicated LLM weights continuously over parallel space-division multiplexed multi-core fibers we bypass the need for massive, redundant localized HBM or DDR memory blocks at every compute node. Our analysis indicates that our design is a reasonable starting point for developing an experimental prototype. By delivering identical weight parameters to thousands of inference pipelines simultaneously, Fiber Memory could cut weightdelivery energy by over 70% compared to HBM3e while reducing data-replication costs. Scaling Fiber Memory to larger models simply requires increasing the number of 50 km fiber segments and PDFAs to increase the amount of data circulating in the loop. Our work should be read as a first-order architectural feasibility study and a call for further study, rather than a finalized implementation blueprint. Future work will investigate the design of weight insertion strategies for growing contexts, advanced synchronization protocols, whether or not the ring network can be used for standard network tasks such as system administration and serving user queries, and custom compiler pipelines that can natively schedule hybrid neural architectures to interlace with light-speed data streams.
References [1] Govind Agrawal. 2012. Fiber-Optic Communication Systems: Fourth Edition. doi:10.1002/9780470918524 [2] Mohamed G. Ahmed, Mohamed Eladawy, John A. Palmer, Ahmed El-Nozahi, and Pavan Kumar Hanumolu. 2021. A 16-Gb/s −11.6-dBm OMA Sensitivity 0.7-pJ/bit Optical Receiver in 65-nm CMOS Enabled by Duobinary Sampling. IEEE Journal of Solid-State Circuits 56, 9 (2021), 2795–2805. doi:10.1109/JSSC.2021.3064248 [3] Reisi Aminabadi, Samyam Rajbhandari, Minjia Zhang, Chuanho Li, et al. 2022. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 255– 267. doi:10.1109/HPCA53960.2022.00027 [4] K. Benyahya, Daniel Burridge, Daniel Cletheroe, Thomas Karagiannis, Brian Robertson, Ant Rowstron, Mengyang Yang, Arash Behziz, Jamie Gaudette, and Paolo Costa. 2025. MOSAIC: Breaking the Optics versus Copper Trade-off with a Wide-and-Slow Architecture and MicroLEDs. In Proceedings of the ACM SIGCOMM 2025 Conference. ACM, Coimbra, Portugal. [5] Brandon Buscaino, Elizabeth Chen, James W. Stewart, Thang Pham, and Joseph M. Kahn. 2021. External vs. Integrated Light Sources for Intra-Data Center Co-Packaged Optical Interfaces. Journal of Lightwave Technology 39 (2021), 1984–1996. doi:10.1109/jlt.2020.3043653 [6] Kun-Zhi Chang, Linhai Zhang, Tiejun Wang, and Intel Communications Group. 2022. Ultra-Low-Latency Hardware Architectures for Reed-Solomon Forward Error Correction in Next-Generation Optical and Electrical Interconnects. IEEE Transactions on Circuits and Systems I: Regular Papers 69, 8 (2022), 3245–3256. doi:10.1109/TCSI.2022. 3168941 7
Hannah Atmer, Thiemo Voigt, Yuan Yao, and Stefanos Kaxiras
[23] NVIDIA Corporation. 2026. GPUDirect. NVIDIA Developer Portal. https://developer.nvidia.com/gpudirect Accessed June 12, 2026. [24] Yasutake Ohishi, Terutoshi Kanamori, Takeshi Kitagawa, Shiro Takahashi, Elias Snitzer, and George H. Sigel. 1998. Praseodymium-Doped Fiber Amplifiers for Optical Communications. IEEE Photonics Technology Letters 10, 4 (1998), 516–518. doi:10.1109/68.662582 [25] Georg Rademacher, Benjamin J. Puttnam, Ruben S. Luis, et al. 2020. High-Capacity Transmission Over Multi-Core Fibers in the O-Band. Journal of Lightwave Technology 38, 2 (2020), 416–422. doi:10.1109/ JLT.2019.2947444 [26] J. Smith. 2025. The DRAM Bottleneck in Generative AI Systems. Forbes Tech Insights (October 2025). [27] Richard Soref, Roberto De Rose, Ran Ding, et al. 2023. Co-Packaged Optics for Next-Generation Data Centers: Power Efficiency and Architecture Analysis. IEEE Communications Magazine 61, 4 (2023), 42–48. doi:10.1109/MCOM.001.2200385 [28] Kristian Stubkjaer, Juerg Leuthold, and Charles H. Joyner. 2019. Power Consumption Dynamics and Thermal Biasing of All-Optical 2R Regenerators Based on Semiconductor Optical Amplifiers. IEEE Photonics Technology Letters 31, 14 (2019), 1143–1146. doi:10.1109/LPT.2019. 2918841 [29] Kazumasa Takada, Takashi Mori, Takemi Hasegawa, and Hanawa Nishi. 2023. Power Consumption Analysis and Gain Optimization of OBand Praseodymium-Doped Fiber Amplifiers for Optical Interconnects. Journal of Lightwave Technology 41, 11 (2023), 3456–3464. doi:10.1109/ JLT.2023.3251104
[30] Min Tan, Jiang Xu, Siyang Liu, et al. 2023. Co-packaged optics (CPO): status, challenges, and solutions. Frontiers of Optoelectronics 16, 1 (2023), 9. doi:10.1007/s12200-023-00058-7 [31] Andrei-Alexandru Ulmămei and Cătălin Bîră. 2026. Reconfigurable SmartNICs: A Comprehensive Review of FPGA Shells and Heterogeneous Offloading Architectures. Applied Sciences 16, 3 (2026), 1476. doi:10.3390/app16031476 [32] Benjamin C. Webb, Tom Farrell, and Robert J. Manning. 2011. AllOptical 2R Regeneration Using Cross-Phase Modulation in a Highly Non-Linear Fiber and an Optical Filter. IEEE Photonics Technology Letters 23, 14 (2011), 983–985. doi:10.1109/LPT.2011.2148107 [33] Jun Shan Wey, Xiang Liu, and Ed Harstead. 2020. Design and Optimization of O-Band DWDM Transmission Systems for High-Speed Mobile Fronthaul and Data Center Networks. Journal of Lightwave Technology 38, 11 (2020), 2913–2921. doi:10.1109/JLT.2020.2978543 [34] Wikipedia contributors. 2026. Delay-line memory — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Delay-line_memory. [Online; accessed 26-June-2026]. [35] Junwen Zhang, Jianjun Yu, Jun Shan Wey, et al. 2020. SOA PreAmplified 100 Gb/s/𝜆 PAM-4 TDM-PON Downstream Transmission Using 10G-Class O-Band Transmitters. Journal of Lightwave Technology 38, 2 (2020), 185–193. doi:10.1109/JLT.2019.2954992 [36] K. Zhang. 2024. Quantifying Cable Lengths and Fiber Densities in Modern Mega-Datacenters. IEEE Communications Magazine 62, 1 (2024), 45–52.
8