1
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
I. I NTRODUCTION Transformer-based networks [1], such as Large Language Models (LLMs) and Vision Transformers (ViTs), have demonstrated state-of-the-art performance (e.g., accuracy) for solving Solomon Micheal Serunjogi and Ayat Taha are with Photonic Research Lab (PRL), Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: [email protected], [email protected]). Rachmad Vidya Wicaksana Putra is with eBRAIN Lab, Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: [email protected]). Muhammad Shafique is the Director of eBRAIN Lab, Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: [email protected]). Mahmoud Rasras is the Director of Photonic Research Lab (PRL), Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates (UAE); (e-mail: [email protected]). ∗ Equal contributions.
Top-1 Accuracy [%]
88 86 84 82 80 78
(b) accuracy improves as the model size increases
ImageNet-1K 0
50
100
λ0 λ1 λ2 Mod Mod
150
200
250
Number of Parameters [M]
300
Mod
Mod
Index Terms—Silicon Photonics, Mode-Division Photonic Accelerator, Transformers, Hardware-Software Co-Design, Inverse Design, Coherent Crossbar.
(a)
Mod
Abstract—Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. However, state-of-the-art rely on expensive multi-wavelength light generation and large dot-product units due to active phase-shifter components, thus making their approach inefficient and impractical. To address this, we propose MDTransformer, a novel hardware-software co-design of PTA based on mode-division optical dataflow and operations. Specifically, MDTransformer performs complex matrix operations using spatial-mode interference, that leverages the inverse-designed multi-mode couplers, crossings, and Mach-Zehnder IQ modulators into a compact mode-division photonic tensor core (MPTC), capable of executing matrix multiplications in the optical domain. Its each guided mode (i.e., TE0 -TE3 ) acts as an independent computational lane, enabling four-fold parallelism-per-waveguide without spectral filtering or free-spectral-range limitations. Moreover, its coherent detection and IQ modulation jointly encode amplitude and phase, realizing complex-valued arithmetic for full-range operations in transformers. MDTransformer offers analog multiplication with sub-4-bit effective precision and intermodal crosstalk below -30 dB. Its inverse-designed approach also offers scalable and full compatibility with single-laser continuouswave operation at 1550 nm. Experimental results show that MDTransformer achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiTTiny/Small/Base and BERT-Base/Large). These results show that MDTransformer offers a practical solution for high-performance and energy-efficient transformer-based systems.
Mod
arXiv:2607.26016v1 [cs.AR] 28 Jul 2026
Solomon Micheal Serunjogi∗ , Rachmad Vidya Wicaksana Putra∗ , Member, IEEE, Ayat Taha, Muhammad Shafique, Senior Member, IEEE, and Mahmoud Rasras, Senior Member, IEEE
λ0 λ1 λ2
DDot
DDot
DDot
DDot
DDot
DDot
DDot
DDot
DDot
Fig. 1. (a) Transformer networks typically improve their performance at the cost of larger memory footprint; based on data from [6]. (b) The state-of-theart photonic tensor core for Transformer acceleration based on dynamicallyoperated dot-product (DDot) unit from the LT accelerator [18].
diverse machine learning (ML) tasks, e.g., natural language processing (NLP) and computer vision [2]–[4], thereby paving the way toward artificial general intelligence (AGI) [5]. This state-of-the-art performance comes at higher computational and memory costs as shown in Fig. 1(a), thereby leading to huge power/energy consumption [6]. This condition limits the wide adoption of Transformer models in diverse application use-cases. Toward this, specialized electronic accelerators for Transformer inference have been developed [7]–[9]. However, such conventional accelerators face challenges as transistor circuits hit the limits of Dennard scaling [10], leading to their diminishing return of performance efficiency (e.g., slower performance gains and increased power dissipation-per-unit area). Recent works have proposed optical-based integrated circuits to expedite neural network (NN) inference, exploiting ultra-high speed and low energy nature of optical-based computation, known as photonic accelerators [11]–[13]. They typically leverage optical components such as Micro-Ring Resonator (MRR) [14] [15], Mach-Zehnder Interferometer (MZI) [16], and Phase Change Material (PCM) [17] for designing a photonic tensor core (PTC). However, these works mainly target convolutional neural networks (CNN) acceleration [18], exposing the need for studies that target transformer acceleration. Therefore, the targeted problem in this work is how can we develop a high performance and energy-efficient photonic transformer accelerator (PTA)? A solution to this problem may enable a practical PTA design for diverse application use-cases. A. State-of-the-Art PTAs and Their Limitations Most of PTA designs employ PTC based on MZI, MRR banks [19]–[23], and PCM crossbars [24]. They statically store operands on the optical components for computation, hence
2
(a) 14% 24.4%
(b)
1.2%
14.8W
26%
12.6%
18.8%
12.4%
26.3%
excluding Micro_comb
2.7%
1.1%
28.1W
25.3%
excluding Micro_comb
20.1%
12.1% 2.9%
Laser
DAC
MZM
ADC
TIA
Core
Adder
Memory
Micro_comb
Fig. 2. Area breakdown of the state-of-the-art 4-bit LT accelerators: (a) LTBase and (b) LT-Large, showing the contributions of different modules.
they suffer from slow operand mapping and programming. Recently, the Lightening-Transformer (LT) accelerator [18] has been proposed. It inspires further studies in reconfigurability aspect [25] and digital-to-analog converter (DAC) optimization [26] [27]. LT improves performance efficiency of Transformer inference over other PTA designs by employing PTC with dynamic operations of full-range input operands through dynamically-operated dot-product (DDot) unit; see Fig. 1(b). Hence, it eliminates slow operand mapping/programming and making it the state-of-the-art PTA design. Despite their benefits, all these works still have the following critical limitations. • They typically employ multiple wavelengths for operations based on wavelength-division multiplexing (WDM), to achieve highly parallel multiply-accumulate (MAC) operations [28] [29]. However, their reliance on finely-spaced resonant filters and dispersion-limited channels imposes scalability bottlenecks, as generating multiple wavelengths consumes huge area and power/energy and is often done in a strongly nonlinear medium such as SiN. • The free-spectral-range (FSR) of MRRs restricts the number of usable wavelengths, while temperature-dependent resonance drift and fabrication non-uniformity demand active thermal control and calibration overhead [30], [31]. • Coherent operation across many wavelength lanes requires precise optical phase alignment and stabilization, thereby adding power and system complexity for a practical solution [32] [33]. • State-of-the-art accelerators relies on Micro comb (MC)based wavelength generator, Mach-Zehnder Modulator (MZM), and phase-shifter (PS)-based PTC, which are area and power hungry. Therefore, they impose scalability and efficiency challenges when designing area- and power-efficient PTA architecture. To show the limitations of state-of-the-art and related research challenges, we conduct a case study in Section I-B. B. Case Study and Related Research Challenges We study the impact of different modules in the state-ofthe-art 4-bit LT accelerators (i.e., LT-Base and LT-Large) [18] on area and power consumption using its open-source codes from the original authors. The experimental results are shown in Fig. 2, from which we draw the following key observations. • Micro comb, MZM, and PS-based PTC jointly occupy 45.5% area in LT-Base and 44.6% area in LT-Large, highlighting their dominant area consumption. • Both LT accelerators incur high power consumption, even without Micro com. It is particularly inefficient for meeting
diverse possible power-constrained computing systems. For instance, embedded AI systems typically require about 5W max. power envelope, which is difficult to meet with the existing solutions. • Employing a Micro comb module, which resides in a separate physical chip, comes with additional complexity and non-trivial challenges since it requires a highly precise control in different aspects, including optical stabilization and power distribution. These observations expose the following key research challenges to address for providing a practical PTA design solution. • The use of Micro comb should be avoided to minimize design complexity and inter-chip communication challenges. Hence, its laser generation functionality should be replaced with an efficient alternative solution. • The area of laser generator, modulator, and PTC should be optimized to significantly reduce area, and hence minimizing power and energy consumption. • PTA architecture and dataflow should be synergistically designed to exploit to maximize the performance and efficiency benefits offered by the optical-based processing. C. Our Novel Contributions To address the targeted problem and related challenges, we propose MDTransformer, a novel hardware-software (HWSW) co-design of photonic transformer accelerator (PTA) that leverages the spatial Mode-Division Multiplexing (MDM), inverse-designed coherent crossbar, and IQ modulation to enable a practical solution for high-performance and energyefficient transformer-based systems. It is also the first work that leverages MDM and inverse design concepts for designing PTA. It employs the following key ideas. • Leveraging Spatial MDM for Photonic Computing (Section III-A). It employs MDM to enable on-chip photonic parallelism by employing efficient Mode-based Multiplexer or Demultiplexer (MUX/DEMUX), which distributes, encodes, and routes a single optical source across multiple guided spatial channels. • MDOT: Mode-Division Dot-Product Unit (Section III-B). It aims to perform dot-product operation that represents multiplication of two full-range operands in multiple modes at sigle frequencies, thereby enabling dynamically-operated processing element (PE) for higher-level architecture hierarchy (i.e., PTC). • MPTC: Mode-Division Photonic Tensor Core (Section III-C). It aims to efficiently accelerate general matrix multiplication (GEMM) by leveraging MDOT, Mode-based MUX/DEMUX, IQ modulator, and crossings in a crossbar array fashion. • Architecture System Design (Section III-C). It aims to develop the architecture system of MDTransformer accelerator, by integrating multiple MPTCs and the supporting digital circuits (e.g., on-chip memory). Furthermore, a dataflow pattern is also developed to maximize the benefits of the MDTransformer architecture. Key Results: We evaluate MDTransformer through functional simulation using Tidy3D [34], as well as hardware eval-
3
Fig. 3. Our proposed MDTransformer Accelerator: (a) MPTC with crossings, modulator, and MDOT; (b) architecture design.
uation (e.g., area, power, energy, and latency) using the stateof-the-art PTA hardware simulator from [18]. Furthermore, our design is also under fabrication. Experimental results show that, MDTransformer offers 4-bit effective precision for multiplication, low inter-modal crosstalk (i.e., below -30 dB), and full compatibility with single-laser continuous-wave operation at 1550nm. MDTransformer also achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiT-T/S/B1 and BERT-B/L2 ). II. P RELIMINARIES A. Transformer-based Network Models A transformer-based network is formed by multiple identical blocks: encoder and decoder blocks. Each block is formed by a multi-head self-attention (MA) module, a feed-forward network (FN) module, a layer normalization (LN) module, and shortcut connections [18]. The decoder block also has cross-attention and masked self-attention modules. The basic encoder block can be stated as Eq. 1-2, where Xl is the input sequences of l-th layer. Multi-head self-attention (MA) module supports H self-attention heads, and each head makes the input vector into separate vectors, i.e., query (Q), key (K), and value (V) vectors. The attention function between these input vectors can be state as Eq. 3, where dk is the dimension of Q and K. Xl+1 = FN(LN(X̂l+1 )) + X̂l+1
(1)
X̂l+1 = MA(LN(Xl )) + Xl QK⊺ Atten(Q, K, V) = sof tmax √ V dk
(2) (3)
B. Inverse Design in Photonic Circuits Inverse-designed components can implement complex transformations, such as filters and couplers, within an order of magnitude smaller than classical designs [35]–[37]. Such 1 DeiT-T/S/B denotes DeiT-Tiny, DeiT-Small, and DeiT-Base, respectively. 2 BERT-B/L denotes BERT-Base and BERT-Large, respectively.
devices maintain low loss and high modal fidelity, making them ideally suited for large-scale photonic accelerators here thousands of operations must be packed into a small area. As photonics moves toward ultra-dense, domain-specific optical computing, inverse design has become a promising approach for building high-performance primitives that enable massive parallelism and energy-efficient linear algebra. III. O UR P ROPOSED MDTransformer ACCELERATOR We develop our PTA design, called MDTransformer, based on MDM, inverse-designed coherent crossbar, and IQ modulation (overview in Fig. 3). Details of its design is discussed in Section III-A - Section III-D. A. Leveraging Spatial Mode-Division Multiplexing for Photonic Computing We leverage spatial MDM to obtain an orthogonal degree of freedom for on-chip photonic parallelism [38], [39], thereby enabling light generation using limited number of on-chip lasers and reducing on the number of WDM channels. Instead of distributing computation across distinct wavelength like in the state-of-the-art works [18], [26], [27], MDM leverages multiple guided modes (e.g., TE0 –TE3 ) within a single multimode waveguide, enabling independent and simultaneous information channels in the spatial domain. Here, operand pairs are mapped onto orthogonal modal channels (e.g., TE0 –TE3 ) of a multi-mode bus waveguide and processed in modular MDOT units. B. MDOT: Mode-Division Dot-Product Unit We propose a Mode-Division Dot-Product Unit (MDOT) to perform a signed multiplication between two full-range modeencoded operands through optical interference; see Fig. 4. Fig. 4(a) shows the novel inverse-designed structure after optimization, which ensures that the coherent coupler occupies a small footprint and does not need an area- and powerhungry 90-degree phase-shifter. Meanwhile, Fig. 4(b) shows the intensity across the MDOT structure. Here, each spatial mode is routed into a compact 8 × 8 µm2 inverse-designed
4
8
0.00
6
-2.00
8000
4.00
7000
2.00
6000 5000
0.00
4000
-2.00
3000
4
-6.00 -6.00 -4.00 -2.00 0.00 2.00 4.00 6.00
2
(m)
IP C ∝ xm ·ym ,
2000
-4.00
-4.00
where xm and ym denote the symbol vectors carried by mode m and C is an arbitrary constant. Summing over all supported spatial modes produces the full mode-division dot product:
6.00
E2
2.00
(b)
𝜀𝑟
10
y (µm)
12
4.00
y (µm)
(a) 6.00
1000
-6.00 -6.00 -4.00 -2.00 0.00 2.00 4.00 6.00
x (µm)
IP C =
x (µm)
Fig. 4. Inverse-designed MDOT design: (a) silicon projection of the design region. (b) Intensity distribution of light traveling through the structure from the left input.
M −1 X m=0
(m)
IP C ∝
M −1 X
xm ·ym .
(7)
m=0
Therefore, the coherent MDOT unit directly performs signed multiplication and accumulation using both optical (high Q resonators) and electrical domain (capacitive dynamics) through time multiplexed integrators [40]–[42]. C. MPTC: Mode-Division Photonic Tensor Core
(a)
(b)
Fig. 5. MDOT properties: (a) simulated relative phase differences between the four output ports; and (b) CMRR.
coherent mixer that produces four output interference states with fixed phase relationships. Each of the four mixer outputs can be described using a linear transformation of the two incoming operands sA and sB ; see Eq. 4. It explicitly shows the {0◦ , 90◦ , 180◦ , −90◦ } phase basis generated at the outputs. [E1 E2 E3 E4 ]⊺ = (4) 1 [(sA + jsB ) (sA − jsB ) (sA − sB ) (sA + sB )]⊺ 2 Figure 5(a) shows the simulated relative phase differences between all four output ports, taken pairwise, when a single mode is launched from the input (left side). Adjacent port pairs (port 0–port 1 and port 2–port 3) maintain approximately 180◦ phase separation across the 1530–1560 nm band, while alternating port pairs maintain roughly 90◦ separation. The figure also shows a near zero port imbalance across the wavelength of interest as shown by the lightly shaded green region. Meanwhile, the corresponding common-mode rejection ratio (CMRR) is shown in Fig. 5(b), exceeding 30 dB at the operating wavelength of λ = 1550 nm. Bipolar Encoding and Signed Multiplication: The MDOT (m) unit accepts two bipolar NRZ symbol streams sx (k) and (m) sy (k) ∈ {−1, +1} derived from input bits ak , bk via s = 1 − 2a. The fields applied to the coherent mixer are: (m)
jϕx Ex(m) (k) = s(m) , x e
(m)
jϕy Ey(m) (k) = s(m) , y e
(5)
with ϕx = πpk and ϕy = πqk and pk , qk ∈ {0, 1} applied through a phase modulator. The balanced detection photocurrent for mode m over N symbol periods yields the mode-wise dot product: Z Nτ (m) (m) IP C = C s(m) x (t) sy (t) cos ∆ϕm (t) dt (6) 0 ∝ xm ·ym ,
We propose a novel photonic tensor-core architecture, referred to as the Mode-Division Photonic Tensor Core (MPTC), for efficient general matrix multiplication (GEMM) using a single optical carrier and multiple orthogonal spatial modes. The MPTC combines four principal building blocks: (i) a mode-based multiplexer/demultiplexer (MUX/DEMUX), (ii) the coherent mode-division dot-product unit (MDOT), (iii) IQbased complex modulation, and (iv) compact routing elements such as crossings and couplers, all assembled into a structured array, as shown in Fig. 3(a). Unlike prior photonic matrix engines that rely on wavelength-division parallelism or cascaded interferometric meshes, the proposed MPTC uses spatial modes as the primary computational lanes. This choice reduces dependence on multiple laser wavelengths, resonance management, and spectral routing overhead, while enabling multiple operands to propagate within the same multimode waveguide. 1) Mode-based Multiplexer and Demultiplexer (MUX and DMUX): The MUX/DEMUX is the front-end modal interface of the MPTC and is responsible for converting single-mode input channels into a multimode computational bus, and conversely extracting specific modes at later processing stages. In contrast to communication-only mode multiplexers, which are typically designed as standalone coupling elements, the MUX/DEMUX here is designed as a computational routing primitive whose role is to inject and recover operands inside a dense coherent dot-product array. Fig. 6 shows the operation of the four-mode MUX/DEMUX, designed using full-wave inverse design in Tidy3D. The figure illustrates the field evolution for each input mode and the corresponding selective routing to the designated output port. Each panel should be interpreted as a mode-resolved demonstration of selective field transformation: for each input channel, the optical energy is redistributed so that only the target output port carries the desired mode, while leakage to the other ports is suppressed. The input interface consists of four single-mode waveguides of width 0.5 µm. This width is selected to ensure robust TE0 operation at λ = 1550 nm, thereby providing a clean modal input state before multiplexing. The spacing between adjacent input waveguides is set to 1.5 µm. This spacing is large enough to suppress unwanted evanescent coupling
5
7.50
y (µm)
2.50
10
0.00
0
-2.50
-10
-5.00
𝑅𝑅𝑅𝑅 𝐸𝐸𝐸𝐸
20
5.00
-20
-7.50 -5.00 -2.50 0.00 2.50 5.00 -5.00 -2.50 0.00 2.50 5.00 -5.00 -2.50 0.00 2.50 5.00 -5.00 -2.50 0.00 2.50 5.00
(a) x (µm)
(b) x (µm)
(c) x (µm)
(d) x (µm)
Fig. 6. Inverse-designed mode-based MUX/DEMUX: optical field distributions for the four input modes, showing selective routing to distinct single-mode outputs. Each panel illustrates a mode-resolved input-to-output field transformation rather than simple power splitting.
between neighboring inputs, yet small enough to maintain dense layout compatibility with the surrounding tensor-core routing network. These four single-mode channels feed a multimode bus waveguide of width 2.5 µm. This width is chosen because it provides a practical trade-off between modal capacity and circuit density: it is sufficiently wide to support the first four guided TE modes (TE0 –TE3 ) at 1550 nm, while remaining narrow enough to avoid excessive crossing area, large bending penalties, and poor array density. In other words, the selected dimensions are not arbitrary; they arise from the joint requirement of supporting four orthogonal modes and embedding them in a compact computational crossbar. To obtain the optimized freeform structure, the permittivity distribution ε(r) is solved through an adjoint-based gradient descent formulation that maximizes the transmission of each target mode into its assigned output port while penalizing leakage into all other ports. The optimization problem is expressed in Eq. 8, where Tm→pm denotes the transmission from input mode m to its designated output port pm , and α is a penalty factor enforcing crosstalk suppression. This objective does not merely maximize throughput; it imposes a modeselective field transformation that preserves the computational meaning of each channel. 3 3 X X max F = Tm→pn Tm→pm − α ε(r)
m=0
(8)
n=0 n̸=m
Physically, the optimized region acts as a compact distributed scattering medium that directly maps one modal basis to another. This is an important distinction from conventional asymmetric directional couplers, microring assisted mode couplers, or MMI-based devices, where coupling is governed by predetermined geometric interference lengths and is often less flexible for simultaneously enforcing compactness, broadband operation, and multi-port modal selectivity. Earlier onchip mode-division multiplexers, such as microring-assisted designs, demonstrated selective mode coupling but remained tied to wavelength-sensitive routing concepts. Similarly, ultracompact multimode routing work has focused on bends and crossings for dense integration. Here, by contrast, the inversedesigned MUX/DEMUX is integrated directly into a photonic
tensor-core data path, where its purpose is not only multiplexing, but controlled operand delivery to coherent compute nodes [39] [43]–[45]. Fabrication-Aware Inverse Design and Constraints: To ensure practical manufacturability, the inverse design process incorporates fabrication-aware constraints consistent with standard electron-beam lithography in silicon photonics. Specifically, a minimum feature size of 120 nm is enforced through spatial filtering and projection steps applied to the permittivity distribution during optimization. This avoids the formation of sub-resolution features and ensures that the final structure can be faithfully fabricated without requiring additional postprocessing. Such feature-size-constrained inverse design has been widely adopted in recent nanophotonic devices to bridge the gap between idealized continuous permittivity optimization and binary fabrication-compatible layouts. Modal Superposition and Decomposition Strategy: Although the device operates on four orthogonal modes simultaneously, the optimization is structured using a modal decomposition approach. Each mode transformation is treated as an independent objective, and the total cost function is constructed as a superposition of these modal targets, as shown in Eq. 8. This ensures that each input mode is selectively mapped to its corresponding output port without interfering with the routing of the remaining modes. This decomposition is physically justified by the orthogonality of the guided modes in the multimode waveguide, allowing independent control of each modal channel while maintaining a shared spatial structure. Reciprocity and Forward–Adjoint Consistency: The optimization leverages electromagnetic reciprocity, whereby the adjoint simulation corresponds to exciting the device from the output ports and propagating fields backward. Consistency between forward and adjoint field distributions ensures that the optimized structure satisfies both excitation and collection conditions simultaneously. In practice, this guarantees that the device performs equivalently under forward multiplexing and reverse demultiplexing operation, which is essential for its dual role within the MPTC architecture. Output Mode Engineering and Power Capture Efficiency: In order to improve power transfer from the multimode region into the output waveguides, the single-mode output
6
6000 5000
E2
4000 3000 2000 1000
20 10 0.0 -10 -20
(a)
(c)
(b)
Re Ez
30
-30
Fig. 7. Cross-sectional optical field distributions for (a) scissors crossing, (b) 50:50 3dB coupler, and (c) 90◦ waveguide crossing.
ports are intentionally widened beyond the nominal 0.5 µm width. Specifically, the outputs are expanded to approximately 0.75 µm before being adiabatically tapered back to standard single-mode dimensions. This local widening improves mode overlap between the transformed field distribution and the guided mode of the output waveguide, thereby enhancing coupling efficiency and reducing scattering loss. Such taperassisted mode matching is critical in inverse-designed structures, where the output field profile may not perfectly match the fundamental mode of a narrow waveguide without additional impedance matching. The field distributions in Fig. 6 also clarify why the inversedesigned approach is needed. For each launched mode, the structure does not simply split power; it redistributes phase and amplitude across a freeform subwavelength region so that the desired output field emerges at one specific port while the remaining ports are suppressed. This mode-resolved routing behavior is exactly what is required in the MPTC: at each downstream computational cell, one selected mode must be exposed to the coherent multiplier, while the remaining modes must continue propagating with minimal disturbance. In benchmarking terms, recent inverse-designed modedivision devices have demonstrated that compact mode multiplexers can significantly outperform conventional moderouting footprints, for example through five-mode inversedesigned MDM devices with a reported footprint of 16×7 µm2 and measured crosstalk below approximately −11 dB, as well as recent scalable mode demultiplexers with sub-1 dB loss and crosstalk below approximately −13 dB at 1550 nm. Dense multimode routing elements such as 8 × 8 µm2 crossings have also been reported for three-mode photonic circuits. Our design inherits the compactness philosophy of these works but targets a different system problem: rather than building a communication link, the present MUX/DEMUX is dimensioned and optimized as the operand-injection and mode-selection interface for a coherent photonic tensor core [39] [45] [46]. Additional supporting components, including the scissors crossing, 50:50 3 dB coupler, and 90◦ waveguide crossing, are presented in Fig. 7(a)–(c), respectively. These elements are used to construct the routing network surrounding the MPTC.
The processing pipeline begins with a continuous-wave input field that is first split using integrated power splitters and 50:50 couplers. These peripheral components also form the basis of other circuit blocks such as IQ modulators and mode-scissors crossings, enabling flexible routing of optical data throughout the processor. After splitting, the optical field is expanded into the multimode bus waveguide and encoded into one of the four orthogonal TE modes used by the MD-Transformer. The inverse-designed MUX/DEMUX then maps each input modal profile onto a unique single-mode output port at 1550 nm. This provides clean modal separation, low inter-mode crosstalk, and a compact footprint suitable for dense dot-product arrays. Overall, the mode-based MUX/DEMUX forms the frontend interface of the MD-Transformer, enabling parallel spatialmode encoding, selective demultiplexing, and physically structured delivery of operands into the downstream coherent processing core. 2) IQ-Based Complex Modulator: To enable full complexvalued encoding of optical operands, each modal channel incorporates a compact IQ-modulator. Two complementary devices are used: (i) a single-input intensity I modulator for amplitude control, and (ii) Q modulator for phase control. Each input channel employs an IQ modulator analogous to that in [47], but modified to operate at high speeds of 25Gb/s. The amplitude branch is defined by the normalized MZI power transfer in Eq. 9, and the quadrature branch applies an additional phase-shift as in Eq. 10. Here, A, B, D, E are empirical calibration coefficients from device-level measurements. 2 PMZI (IA ) = 12 + 12 cos AIA + BIA + ϕA , (9) 2 ϕPS (IQ ) = DIQ + EIQ + ϕQ ,
(10)
Combining these responses, the complex field at the modulator output for mode m is defined as: p (m) Ex(m) (IA , IQ ) = |E0 | PMZI (IA ) exp i ϕPS (IQ ) . (11) D. Architecture System Design 1) Overall System: The architectural system of our proposed MDTransformer accelerator is shown in Fig. 3(b). A
7
B
ADC Accum.
Tile 0
mode3 mode2 mode1 mode0
mode0
…
Σ
crossing
mode3 mode2 mode1
=
mode3 mode2 mode1 mode0 x3 x2 x1 x0
… Buffer
M2
× ×
y3 y2 y1 y0 mode3 mode2 mode1 mode0 x0 . y0 x1 . y1 time
…
Dh
×
out 1
× ×
(b)
… … …
MPTC
M1
Tile 1
out 0
MPTC
Dv
Tile 0
MPTC
A
×
MPTC
(a) a portion of data
MDOT
Fig. 8. Dataflow based on data tiling mechanism for MDTransformer. (a) It partitions data from M1 across Dv , and map them across tiles. (b) Its crossing routes operands based on their mode to the corresponding MDOT.
2) Dataflow: To maximize benefits of the MDTransformer architecture, a specialized dataflow is developed. Its key ideas are illustrated in Fig. 8 and described below. Multiple data is processed in the same MDOT without any prior data programming; see Fig. 8(a). Multiple tiles can process multiple portions of data, which determines the parallelism level in a chip; see A . Then, multiple cores (MPTCs) can process a portion of data, which also determines the parallelism level in a tile; see B . Afterward, an Nh ×Nv MDOT array can perform multiplications in parallel in the core level. • Each MDOT performs multiplication between two operands from the same mode. Hence, a sequence of multiplications can be scheduled to be performed in the same MDOT, enabling flexible scheduling for exploiting data reuse without expensive broadcast routing; see Fig. 8(b).
MDTransformer employs Nt =4 tiles, Nc =2 cores-per-tile, Nh =Nv =4 input horizontal/vertical waveguides-per-core, Nλ =1 wavelength, and 4 modes. • LT-Base employs Nt =4, Nc =2, and Nh =Nv =Nλ =12. • LT-Large employs Nt =8, Nc =2, and Nh =Nv =Nλ =12. • LT-Custom employs Nt =4, Nc =2, and Nh =Nv =Nλ =4. LT-Large employs 4MB global SRAM, while the others use 2MB. Our design is under fabrication and its measurement setup is shown in Fig. 9(b). •
MDTransformer Arch. Configuration
Workloads
(i.e., DeiT-T/S/B and BERT-B/L)
(e.g., Nt, Nc, Nh, Nv)
Photonic Devices
(Modulator, etc.)
& Circuits
(Crossing, etc.)
•
(a)
PTA Hardware Simulator (.py)
Photonic Simulator (Tidy3D)
Functionality & Characteristics Results
Performance & Efficiency Results
(e.g., CMRR, Imbalance)
(i.e., area, power, energy, latency)
RF Probes
Measurement Setup
single MDTransformer chip has Nt tiles, and each tile consists of Nc MPTCs. An MPTC contains an array of Nh ×Nv MDOTs. Furthermore, MDTransformer also employs on-chip global SRAM whose size should be at least meeting the minimum required size for storing the largest activations in a layer; following the LT design [18]. The global SRAM size should not be significantly smaller or larger than this minimum required size, because it can increase the costly off-chip data access (i.e., high access latency and energy) or aggravate the static power consumption, respectively [48]–[50].
Optical Fiber
Optical DC Probes Chip
(b)
Fig. 9. (a) Experimental setup and tools flow in this work. (b) Measurement setup for testing the fabricated chip. TABLE I S UMMARY OF DEVICE PARAMETERS Device DAC [51]
ADC [52] TIA [53]
IV. E VALUATION M ETHODOLOGY
MZM
To evaluate our MDTransformer design, we employ: (1) functional simulation using Tidy3D [34], and (2) hardware evaluation using the state-of-the-art PTA hardware simulator from [18]; see Fig. 9(a). We use functional simulation to evaluate the functionality and characteristics of our proposed optical devices and circuits. The corresponding results are mainly presented in Section III to validate the functionality of MDTransformer. Meanwhile, we use PTA hardware simulator aims to evaluate area and power of the design as well as its energy consumption and latency when running the workload, while considering device parameters from measurements; see Table I. We select DeiT-T, DeiT-S, DeiT-B, BERT-B, and BERT-L as the workloads. As comparison partners, we use the state-of-the-art LT-Base, LT-Large, and LT-Custom with the following configurations.
Crossing Phase Shifter Y-Splitter 4 Mode MUX/DEMUX Coherent Hybrid Photodetector
Parameter Precision Power Area Precision Power Area Power Area Power Area IL Area IL Area IL Area IL Area IL Area Power Sensitivity Area
Value 8-bit 42 mW (@28 GSPS) 0.03 mm2 32 lines, 6-bit 410 mW (@12.8 GSPS) 780 µm2 30 mW < 50 µm2 50 mW ∼2,260 µm2 0.3 dB 36 µm2 1 dB 250 µm2 0.4 dB 36 µm2 6.2 dB 64 µm2 6 dB 144 µm2 1.1 mW -25 dBm 4×10 µm2
V. R ESULTS AND D ISCUSSION A. Reduction of Area and Power Consumption Experimental results for area and power consumption are presented in Fig. 10(a)-(d). The results show that MDTrans-
8
Modulator ADC
697.9 mW
TIA PD
1 88.4%
Energy [mJ]
1 0.8 0.6
DeiT-T
0.4
Adder Memory
0.4%5.9% 0.9% 0.8 4
(e.1)
4 6
100 80 60 40 20 0
DAC
12.8% 0.6%
16.6 mm2
(c) 120
Laser
0.1%
5
0.6 3
(e.2)
4
DeiT-S
5
6
0.4 2
0
0 Embed
QKV
Atten
2.0 16 1.5 12 1.0 8
LT-Large LT-Base LT-Custom
(e.3)
5
DeiT-B
4
6
0.5 4
0.2 1
0.2
Significant area savings compared to other state-of-the-art designs
Proj
FFN1
FFN2
Head
6 4 2
0.0 0
0
8
0 Others
Latency
(d) 30 2
MD Transformer
10 8 6
25 20 15 10 5 0
Significant power savings compared to other state-of-the-art designs
LT-Large
(e.4)
6.0 80
5
6
4
3
LT-Base LT-Custom MD Transformer
4
BERT-B
Micro_comb Memory Adder Core TIA ADC Modulator DAC Laser
4.5 60
1.5 20
0
0.0 0
1
2
3
4
QKV
Atten
Proj
FFN1
FFN2
4
BERT-L
5
6
3.0 40
2
60
(e.5)
45 30 15
1 Head
2
3 Others
4
Latency [ms]
(b) Power Breakdown
Power [W] Thousands
4.3% 2.6% 0.5% 1.1% 2.8%
Area [mm2]
(a) Area Breakdown
0
Latency
Fig. 10. Experimental results for (a) area breakdown of MDTransformer; (b) power breakdown of MDTransformer; (c) comparison on area; (d) comparison on power; as well as energy consumption and latency for different workloads: (e.1) DeiT-T, (e.2) DeiT-S, (e.3) DeiT-B; (e.4) BERT-B, and (e.5) BERT-L.
former occupies 16.6mm2 area and incurs 697.9mW power; see 1 . These profiles are dominated by on-chip memory as the impact of Micro comb, modulator, and phase-shifter is significantly decreased compared to state-of-the-art designs. The reason is that, our design strategy for developing MDTransformer is to eliminate Micro comb, reduce modulator size, and remove phase-shifter in MDOT, hence leading to significantly small area and low power consumption. Area comparison: Our MDTransformer significantly saves area compared to all state-of-the-art designs, i.e., reducing area by 85.3% from LT-Large, 72.4% from LT-Base, and 40.3% from LT-Custom; see 2 . MDTransformer occupies smaller area than LT-Custom despite having the same number of tiles, cores, and core size. The reason is that, MDTransformer employs smaller modulator as well as eliminates Micro comb and phase-shifter in MDOT. In addition to that, MDTransformer also employs smaller number of tiles and smaller core size compared to LT-Large and LT-Base, thus leading to significantly smaller area. Power comparison: Our MDTransformer significantly decreases power compared to all state-of-the-art designs, i.e., reducing power consumption by 97.5% from LT-Large, 95.3% from LT-Base, and 63.6% from LT-Custom; see 3 . MDTransformer incurs smaller power than LT-Custom despite having the same number of tiles, cores, and core size. The reason is that, MDTransformer employs efficient modulator design as well as completely removes power consumption from Micro comb and phase-shifter in MDOT. In addition to that, MDTransformer also employs smaller number of tiles and smaller core size compared to LT-Large and LT-Base, thus leading to significantly lower power consumption. B. Enabling High Performance and Energy Efficiency across Transformer Workloads Experimental results for energy consumption and latency of core processing on different workloads are provided in Fig. 10(e.1)-(e.5). Based on these results, we make the following key observations.
MDTransformer consistently achieves lower energy consumption than LT-Custom across different workloads, despite having the same number of tiles, cores, and core size; as indicated by 4 . Specifically, MDTransformer saves energy consumption by 43.1%-43.5% for DeiT-based models and by 40.6%-45.1% for BERT-based models as compared to LT-Custom. The reason is that, MDTransformer eliminates power requirement for phase-shifter in MDOT and reduces power cost for modulator, which in turns leading to lower energy consumption when processing the workload. • MDTransformer achieves comparable processing latency to LT-Custom across different workloads, as shown by 5 . The reason is that, these two designs consider the same onchip memory size as well as the same number of tiles, cores, and core size. Therefore, they have similar capabilities in storing data on-chip and performing computation based on their dataflow and scheduling. However, such a similar performance comes at the different cost of power consumption, as shown by 3 . Therefore, their energy consumption profiles also differ significantly across different workloads, as indicated by 4 . LT-Large and LT-Base consume relatively low energy since they employ high parallelism to expedite the processing, hence leading to low latency. However, this comes at the cost of huge power consumption, as indicated in Fig. 10(d). This condition may limit the applicability of the LT-Large and LT-Base accelerators for diverse lowpower application use-cases. In contrast, MDTransformer achieves competitive energy consumption compared to LTLarge and LT-Base, as shown by 6 , while incurring a significantly lower power consumption than LT-Large and LT-Base, as shown by 3 . The reason is that, MDTransformer combines the benefits of low-power design through optimized optical devices/modules, selection of architecture configuration, and efficient dataflow for enabling highperformance and energy-efficient optical-based processing. •
C. Further Discussion In this work, we consider a configuration of Nt =4, Nc =2, Nh =Nv =4, Nλ =1, and 4 modes for our MDTransformer.
9
However, this selection of configuration can be adjusted based on the requirements. For instance, if we need to increase the parallelism in the MDTransformer, then we can increase the number of tiles Nt , number of cores-per-tile Nc , and core size Nh xNv . Conversely, if we have a targeted application that imposes tight design constraints, e.g., in terms of area, power, energy, and latency (or throughput), then the configuration should be selected carefully. All these adjustment choices are supported with our dynamically-operated architecture and dataflow design in MDTransformer, thereby providing a practical, high-performance and energy-efficient PTA design. VI. C ONCLUSION We propose a novel hardware-software co-design of MDTransformer accelerator, which employs MDM-based computation, inverse-designed coherent PTC, and IQ modulation. Experimental results show that, our MDTransformer accelerator offers 4-bit effective precision for multiplication, low intermodal crosstalk (i.e., less than -30dB), and full compatibility with single-laser continuous-wave operation at 1550nm. It also saves 40.4% area, 63.6% power, and 40.6% energy consumption, with comparable latency over the state-of-the-art across different transformer models. Therefore, our MDTransformer accelerator successfully provides a practical solution for highperformance and energy-efficient transformer-based systems. R EFERENCES [1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez et al., “Attention is all you need,” Advances in Neural Information Processing Systems (NIPS), vol. 30, no. 1, pp. 261–272, 2017. [2] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021. [3] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML). PMLR, 2021, pp. 10 347–10 357. [4] S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM Computing Surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022. [5] G. Yenduri, R. Murugan, P. Kumar Reddy Maddikunta, S. Bhattacharya, D. Sudheer, and B. Bhushan Savarala, “Artificial general intelligence: Advancements, challenges, and future directions in agi research,” IEEE Access, vol. 13, pp. 134 325–134 356, 2025. [6] K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao, “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 45, no. 1, pp. 87–110, 2023. [7] H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021, pp. 97–110. [8] M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memorybased acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, pp. 1071–1085. [9] H. You, Z. Sun, H. Shi, Z. Yu, Y. Zhao, Y. Zhang, C. Li, B. Li, and Y. Lin, “Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2023, pp. 273–286. [10] F. P. Sunny, E. Taheri, M. Nikdast, and S. Pasricha, “A survey on silicon photonics for deep learning,” ACM Journal of Emerging Technologies in Computing System (JETC), vol. 17, no. 4, pp. 1–57, 2021.
[11] K. Shiflett, A. Karanth, R. Bunescu, and A. Louri, “Albireo: Energyefficient acceleration of convolutional neural networks via silicon photonics,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2021, pp. 860–873. [12] B. J. Shastri et al., “Photonics for artificial intelligence and neuromorphic computing,” Nature Photonics, vol. 15, no. 2, 2021. [13] Z. Yin, M. Zhang, N. Gangi, R. Huang, J. Zhang, and J. Gu, “Simphony: A device-circuit-architecture cross-layer modeling and simulation framework for heterogeneous electronic-photonic ai system,” in 2025 62nd ACM/IEEE Design Automation Conference (DAC). IEEE, 2025, pp. 1–7. [14] A. N. Tait, T. F. De Lima, E. Zhou, A. X. Wu, M. A. Nahmias, B. J. Shastri, and P. R. Prucnal, “Neuromorphic photonic networks using silicon photonic weight banks,” Scientific Reports, vol. 7, no. 1, p. 7430, 2017. [15] F. Sunny, A. Mirza, M. Nikdast, and S. Pasricha, “Crosslight: A crosslayer optimized silicon photonic neural network accelerator,” in 2021 58th ACM/IEEE design automation conference (DAC). IEEE, 2021, pp. 1069–1074. [16] Y. Shen, N. C. Harris, S. Skirlo, M. Prabhu, T. Baehr-Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englund et al., “Deep learning with coherent nanophotonic circuits,” Nature photonics, vol. 11, no. 7, pp. 441–446, 2017. [17] J. Feldmann, N. Youngblood, M. Karpov, H. Gehring, X. Li, M. Stappers, M. Le Gallo, X. Fu, A. Lukashchuk, A. S. Raja et al., “Parallel convolutional processing using an integrated photonic tensor core,” Nature, vol. 589, no. 7840, pp. 52–58, 2021. [18] H. Zhu, J. Gu, H. Wang, Z. Jiang, Z. Zhang, R. Tang, C. Feng, S. Han, R. T. Chen, and D. Z. Pan, “Lightening-transformer: A dynamicallyoperated optically-interconnected photonic transformer accelerator,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2024, pp. 686–703. [19] Y. Li, A. Louri, and A. Karanth, “Sprint: A high-performance, energyefficient, and scalable chiplet-based accelerator with photonic interconnects for cnn inference,” IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 10, pp. 2332–2345, 2022. [20] ——, “Spacx: Silicon photonics-based scalable chiplet accelerator for dnn inference,” in 2022 IEEE International Symposium on HighPerformance Computer Architecture (HPCA), 2022, pp. 831–845. [21] S. Afifi, F. Sunny, M. Nikdast, and S. Pasricha, “Tron: Transformer neural network acceleration with non-coherent silicon photonics,” in Great Lakes Symposium on VLSI (GSVLSI) 2023, 2023, pp. 15–21. [22] S. Afifi, O. Alo, I. Thakkar, and S. Pasricha, “A light-speed large language model accelerator with optical stochastic computing,” in Great Lakes Symposium on VLSI (GLSVLSI) 2025, 2025. [23] ——, “Astra: A stochastic transformer neural network accelerator with silicon photonics,” ACM Transactions on Embedded Computing Systems (TECS), 2025. [24] Y. Li, A. Louri, and A. Karanth, “Merit: A sustainable dnn accelerator design with photonic phase-change memory,” IEEE Transactions on Sustainable Computing (TSUSC), vol. 10, no. 4, pp. 705–716, 2025. [25] H. Zhu, Z. Zhou, S. Ning, X. Wu, R. Chen, Y. Wan, and D. Pan, “Enlighten: Lighten the transformer, enable efficient optical acceleration,” arXiv preprint arXiv:2510.01673, 2025. [26] H. Li, D. Chen, and T. Mitra, “Hyatten: Hybrid photonic-digital architecture for accelerating attention mechanism,” in 2025 Design, Automation & Test in Europe Conference (DATE), 2025, pp. 1–7. [27] W.-T. Chang, C.-F. Wu, and Y.-C. Lo, “P-dac: Power-efficient photonic accelerators for llm inference,” in 2025 62nd ACM/IEEE Design Automation Conference (DAC), 2025, pp. 1–7. [28] R. Hamerly, A. Sludds, S. Bandyopadhyay, Z. Chen, Z. Zhong, L. Bernstein, and D. Englund, “Netcast: low-power edge computing with wdmdefined optical neural networks,” Journal of Lightwave Technology, vol. 42, no. 22, pp. 7795–7806, 2024. [29] H. Li, D. Chen, and T. Mitra, “Hybrid photonic-digital accelerator for attention mechanism,” arXiv preprint arXiv:2501.11286, 2025. [30] S. Biasi, G. Donati, A. Lugnan, M. Mancinelli, E. Staffoli, and L. Pavesi, “Photonic neural networks based on integrated silicon microresonators,” Intelligent Computing, vol. 3, p. 0067, 2024. [31] A. N. Tait, A. X. Wu, T. F. De Lima, E. Zhou, B. J. Shastri, M. A. Nahmias, and P. R. Prucnal, “Microring weight banks,” IEEE Journal of Selected Topics in Quantum Electronics, vol. 22, no. 6, pp. 312–325, 2016. [32] S. Banerjee, M. Nikdast, and K. Chakrabarty, “Characterizing coherent integrated photonic neural networks under imperfections,” Journal of lightwave technology, vol. 41, no. 5, pp. 1464–1479, 2022.
10
[33] A. Totovic, G. Giamougiannis, A. Tsakyridis, D. Lazovsky, and N. Pleros, “Programmable photonic neural networks combining wdm with coherent linear optics,” Scientific reports, vol. 12, no. 1, p. 5605, 2022. [34] I. Flexcompute, “Tidy3D: Next-generation electromagnetic simulation tool,” https://www.flexcompute.com/tidy3d/solver/, 2024, accessed: 2025-01-01. [35] N. V. e. a. Sapra, “Inverse design of compact multimode multi-port photonic devices,” Nature Communications, vol. 11, p. 6361, 2020. [36] J. S. e. a. Jensen, “Adjoint-based inverse design of efficient, broadband mode conversion devices,” ACS Photonics, vol. 7, pp. 1497–1506, 2020. [37] D. e. a. Vercruysse, “Compact broadband directional couplers using inverse design,” Optica, vol. 7, pp. 179–185, 2020. [38] Y. Wang, Y. Wei, V. Dolores-Calzadilla, K. Williams, M. Smit, D. Dai, and Y. Jiao, “Mode division multiplexing on an inp membrane on silicon,” Optics Letters, vol. 47, no. 16, pp. 4004–4007, 2022. [39] Y. Liu, K. Xu, S. Wang, W. Shen, H. Xie, Y. Wang, S. Xiao, Y. Yao, J. Du, Z. He et al., “Arbitrarily routed mode-division multiplexed photonic circuits for dense integration,” Nature communications, vol. 10, no. 1, p. 3263, 2019. [40] S. Lam, A. Khaled, S. Bilodeau, B. A. Marquez, P. R. Prucnal, L. Chrostowski, B. J. Shastri, and S. Shekhar, “Dynamic electro-optic analog memory for neuromorphic photonic computing,” arXiv preprint arXiv:2401.16515, 2024. [41] S. Ning, H. Zhu, C. Feng, J. Gu, Z. Jiang, Z. Ying, J. Midkiff, S. Jain, M. H. Hlaing, D. Z. Pan et al., “Photonic-electronic integrated circuits for high-performance computing and ai accelerators,” Journal of Lightwave Technology, 2024. [42] H. Babashah, Z. Kavehvash, A. Khavasi, and S. Koohi, “Temporal analog optical computing using an on-chip fully reconfigurable photonic signal processor,” Optics & Laser Technology, vol. 111, pp. 66–74, 2019. [43] L.-W. Luo, N. Ophir, C. P. Chen, L. H. Gabrielli, C. B. Poitras, K. Bergmen, and M. Lipson, “Wdm-compatible mode-division multiplexing on a silicon chip,” Nature communications, vol. 5, no. 1, p. 3069, 2014. [44] K. Y. Yang, C. Shirpurkar, A. D. White, J. Zang, L. Chang, F. Ashtiani, M. A. Guidry, D. M. Lukin, S. V. Pericherla, J. Yang et al., “Multidimensional data transmission using inverse-designed silicon photonics and microcombs,” Nature communications, vol. 13, no. 1, p. 7862, 2022. [45] J. L. Pita Ruiz, N. Dalvand, and M. Ménard, “Integrated silicon nitride devices via inverse design,” Nature Communications, vol. 16, no. 1, p. 9307, 2025. [46] J. Li, X. Li, L. Wu, M. Luo, Y. Li, Y. Wang, and Y. Qiu, “Ultra-compact scalable mode demultiplexers for high-speed optical interconnects via gpu-accelerated inverse design,” Optics Express, vol. 33, no. 21, pp. 44 908–44 924, 2025. [47] S. Rahimi Kari, N. A. Nobile, D. Pantin, V. Shah, and N. Youngblood, “Realization of an integrated coherent photonic platform for scalable matrix operations,” Optica, vol. 11, no. 4, pp. 542–551, 2024. [48] R. V. W. Putra, M. A. Hanif, and M. Shafique, “Drmap: A generic dram data mapping policy for energy-efficient processing of convolutional neural networks,” in 2020 57th ACM/IEEE Design Automation Conference (DAC), 2020, pp. 1–6. [49] ——, “Romanet: Fine-grained reuse-driven off-chip memory access management and data organization for deep neural network accelerators,” IEEE Transactions on Very Large Scale Integration Systems (TVLSI), vol. 29, no. 4, pp. 702–715, 2021. [50] ——, “Pendram: Enabling high-performance and energy-efficient processing of deep neural networks through a generalized dram data mapping policy,” arXiv preprint arXiv:2408.02412, 2024. [51] P. Caragiulo, O. E. Mattia, A. Arbabian, and B. Murmann, “A 2x timeinterleaved 28-gs/s 8-bit 0.03-mm 2 switched-capacitor dac in 16-nm finfet cmos,” IEEE Journal of Solid-State Circuits, vol. 56, no. 8, pp. 2335–2346, 2021. [52] Y. Duan and E. Alon, “A 12.8 gs/s time-interleaved adc with 25 ghz effective resolution bandwidth and 4.6 enob,” IEEE Journal of SolidState Circuits, vol. 49, no. 8, pp. 1725–1738, 2014. [53] S. Serunjogi, M. Rasras, and M. Sanduleanu, “64gb/s nrz/pam4 burstmode optical receiver frontend with gain control, offset correction and gain decoupled from bandwidth,” in 2021 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 2021, pp. 1–4.