ConceptioArchivearXiv CS
arXiv CSopen access

DxPTA: An Architecture Design Space Exploration with Optical Dataflow-guided Strategy for HW/SW Co-Design of Photonic Transformer Accelerators

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

1

DxPTA: An Architecture Design Space Exploration with Optical Dataflow-guided Strategy for HW/SW Co-Design of Photonic Transformer Accelerators

arXiv:2606.06515v1 [cs.AR] 2 Jun 2026

Rachmad Vidya Wicaksana Putra, Member, IEEE, Solomon Micheal Serunjogi, Mahmoud Rasras, Senior Member, IEEE, and Muhammad Shafique, Senior Member, IEEE

Abstract—Transformer-based networks have emerged as prominent AI models with state-of-the-art performance, which potentially pave the way toward artificial general intelligence (AGI). However, their large sizes still hinder their efficient implementation, thus highlighting the need for alternate solutions to enable their energy-efficient acceleration. Recently, state-ofthe-art works propose photonic transformer accelerators (PTAs) with significant speedup and energy efficiency improvements over the conventional electronic accelerators. However, their PTA architectures are developed without considering the application constraints (e.g., area, power, energy, and latency). Moreover, their manual design approach also requires huge design time to determine a suitable architecture for the targeted application, hence making this approach not scalable. To address these limitations, we propose DxPTA, a novel design space exploration methodology for enabling efficient hardware/software co-design of the appropriate PTA architecture that meets all constraints. It is achieved by (1) identifying the PTA architecture parameters based on the coherent optical dataflow; (2) analyzing the impact/significance of the parameters; and (3) leveraging this analysis for devising a constraint-aware architecture search algorithm. Experimental results show that, our DxPTA can find the appropriate PTA architectures for different transformerbased models (i.e., DeiT-T/S/B and BERT-B/L). It achieves up to 26mm2 area, 4.8W power, 39mJ energy, and 6ms latency, for constraints of 50mm2 area, 5W power, 50mJ energy, and 10ms latency; with 15.2x faster searching time than the exhaustive approach. These results demonstrate the potential of DxPTA methodology for enabling efficient PTA designs for diverse AGIbased applications. Index Terms—Silicon Photonics, Photonic Transformer Accelerator (PTA), Design Space Exploration (DSE), Coherent Optical Dataflow, Hardware/Software (HW/SW) Co-Design.

I. I NTRODUCTION Transformer-based network models [1], such as Vision Transformers (ViTs) and Large Language Models (LLMs), have emerged as prominent AI models with state-of-the-art performance for solving diverse machine learning tasks, such Rachmad Vidya Wicaksana Putra is with eBRAIN Lab, Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: [email protected]). Solomon Micheal Serunjogi is with Photonic Research Lab (PRL), Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: [email protected]). Mahmoud Rasras is the Director of Photonic Research Lab (PRL), Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates (UAE); (e-mail: [email protected]). Muhammad Shafique is the Director of eBRAIN Lab, Division of Engineering, New York University (NYU) Abu Dhabi, United Arab Emirates; (e-mail: [email protected]).

as vision and natural language processing (NLP) [2]–[5], hence potentially paving the way toward artificial general intelligence (AGI) [6] [7]. However, this state-of-the-art performance comes at high computational and memory requirements as shown in Fig. 1(a), hence leading to huge power/energy consumption [5]. This condition makes it difficult to obtain high performance efficiency when processing transformer models for wide-scale implementations across diverse applications. A potential solution is employing specialized accelerators for expediting the inference of transformers, thus minimizing power consumption and improving energy efficiency [8]–[14]; see Fig. 1(b). However, conventional electronic accelerators face challenges related to their diminishing performance efficiency (i.e., increased power dissipation-per-unit area and slower performance gains), as transistor circuits reach the limits of Dennard scaling [15]. Recently, electronic-photonic integrated circuit (EPIC)based solutions, so-called photonic accelerators [16]–[19], have been studied as an alternative for achieving significant speedup and efficiency improvements over the electronic accelerators, due to their ultra-high speed, high bandwidth, and low energy consumption [18]. Therefore, employments of photonic accelerators for expediting neural network (NN) workloads are actively being explored. For instance, photonic tensor core (PTC) developments leveraging optical components such as Mach-Zehnder Interferometer (MZI) [20], Micro-Ring Resonator (MRR)-based bank [21] [22], and Phase Change Material (PCM)-based crossbar [23]. However, these works still target in accelerating traditional convolutional neural networks (CNN) workloads [24], hence indicating the need for further studies to enable high-performance and energy-efficient inference of transformer models. Therefore, the targeted research problem in this paper is how can we effectively enable high performance and energy-efficient inference of transformer models using photonic-based accelerators? A solution to this problem may enable the efficient deployments of transformerbased models on photonic-based computing systems. A. State-of-the-art of Photonic Transformer Accelerators (PTAs) and Their Limitations Recent works propose PTA designs that employ staticallyoperated PTCs using MRR banks [27]–[31] and PCM crossbars [32]. Another work proposes the Lightening-Transformer (LT) [24] based on dynamically-operated PTCs. It inspires further studies in digital-to-analog converter (DAC) design [33],

2

Transformer-based models typically achieve higher accuracy at the cost of larger memory footprints

84

increasing accuracy as the model size increases

82 80 78

(a)

0

100

200

300

Number of Parameters [M]

(b)

Fig. 1. (a) Transformer-based models typically can improve the performance at the cost of larger memory size (i.e., higher number of parameters); based on the data from [5]. (b) Experimental results of running Data-efficient Image Transformer Base (DeiT-B) [3] with different compute platforms: CPU, GPU, CMOS-based accelerators (i.e., AutoViT-4bit [25] and HeatViT-8bit [26]) and photonic-based Lightening-Transformer (LT) accelerators (i.e., LT-Base-4bit and LT-Large-4bit); based on data from [24].

A

10 0

2

Thousands Power [W]

60 45 30 15 0

4 6 8 Number of Tiles (Nt) Nt=4 & Nc varies

C A 2 4 6 8 Number of Cores-per-Tile (Nc)

80 40 0

8

240 180 120 60 0

Nt varies & Nc = 2

6 4

B

2 0 8

2

4 6 8 Number of Tiles (Nt) Nt=4 & Nc varies

6 4

B

2 0

2 4 6 8 Number of Cores-per-Tile (Nc)

0.8

Latency [ms]

20

120

0.6 0.4 0.2 0

0.4

Latency [ms]

C

Energy [mJ]

30

(b) Energy Consumption & Latency

160

Energy [mJ]

Thousands Power [W]

Nt varies & Nc = 2

Area [mm2]

40

Area [mm2]

(a) Power Consumption & Area

Different configurations lead to different profiles of area, power, energy, and latency, thus highlighting the wide range of design choices for developing photonic accelerators. • Increasing Nt or Nc leads to higher power and larger area due to more complex circuitry (see A ), but it may reduce latency and energy consumption due to increased parallelism (see B ). This shows the need for trade-off analysis in accelerator design. • The state-of-the-art LT design (with Nt =4 and Nc =2) may not meet the constraints. For instance, in low-power applications with max. 5W, the design incurs significantly more power (∼15W) as shown by C , indicating the need for a custom architecture. These observations expose several research challenges in devising solutions for the targeted research problem, as outlined below. • The solution should leverage the characteristics of photonic devices and optical dataflow to find the PTA architecture that meets all constraints, ensuring its applicability for diverse applications. • The solution should minimize the searching time of PTA architecture, hence expediting the design time and providing a scalable design approach for diverse applications. •

Energy [mJ] (log scale)

ImageNet-1K

86

DeiT-B-224

Throughput [FPS] (log scale)

Top-1 Accuracy [%]

88

10000 1000 100 10 1 10000 1000 100 10 1

0.3 0.2 0.1 0

Fig. 2. Experimental results considering different configurations of architecture parameters (i.e., Nt and Nc ) in the 4-bit LT accelerator for (a) power and area; and (b) energy consumption and latency considering the DeiT-Base [3].

[34] and reconfigurability [35]. LT improves the performance and efficiency of transformer inference over other PTAs by enabling dynamic operations of full-range input operands, making it the state-of-the-art PTA design. Despite their benefits, all these works still have the following limitations. Their architectures are developed without considering application constraints (e.g., area, power, energy, and latency), and hence their designs are not directly applicable for targeted applications and leading to sub-optimal performance and efficiency gains. • Their manual design approach needs a huge design time and power/energy consumption to develop a suitable architecture for the targeted applications, thus making this approach not scalable. •

To illustrate the limitations of state-of-the-arts and related research challenges, we perform an experimental case study, which will be discussed in Section I-B.

B. Case Study and Research Challenges We explore the impact of different architecture parameters of the state-of-the-art 4-bit LT accelerator [24]. Here, we vary the number of tiles (Nt ) and the number of cores-per-tile (Nc ). For workload, we consider the DeiT-Base model [3]. Details of the LT hardware architecture and the experimental setup are provided in Section II-B and Section IV, respectively. Experimental results are shown in Fig. 2, from which we make the following key observations.

C. Our Novel Contributions To address the targeted research problem and related challenges, we propose DxPTA, a novel architecture Design space exploration methodology leveraging coherent optical dataflowguided strategy for efficient hardware/software (HW/SW) codesign of Photonic Transformer Accelerators while meeting multiple constraints (i.e., area, power, energy, and latency). It employs the following key steps (see an overview in Fig. 3 and details in Fig. 4). • Identify the architecture parameters of PTA (Section III-A): It targets to analyze the characteristics of PTA architecture (including its hierarchy and photonic devices) and its coherent optical dataflow for identifying the prominent architecture parameters. • Analyze the impact of architecture parameters (Section III-B): It identifies the significance of architecture parameters (e.g., Nt and Nc ) by observing their impact on area, power, energy, and latency. The information will be leveraged for architecture search. • Devise the constraint-aware search algorithm (Section III-C): It explores architecture candidates by leveraging the coherent optical dataflow and parameter significance, evaluates their energy-delay products (EDPs), and then select the one that has the lowest EDP and meets all constraints. Key Results: We evaluate our DxPTA methodology using Python implementation, and then run it on an Nvidia RTX 6000 Ada GPU machine, while considering diverse transformer workloads (i.e., DeiT-T/S/B1 and BERT-B/L2 ). Experimental results show that, DxPTA successfully finds the 1 DeiT-T/S/B refers to DeiT-Tiny, DeiT-Small, and DeiT-Base, respectively. 2 BERT-B/L refers to BERT-Base and BERT-Large, respectively.

3

tile

core

PE

… core

core

core

core

PE

PE

PE

Photonic Device & Circuit Models

PTA Architecture with Selected Configuration

Analyze the Impact of Architecture Parameters (Section III-B)

core

core

core

PE

core

Devise the Constraint-aware Search Algorithm (Section III-C)

(i.e., area, power, energy, and latency)

Workloads

(i.e., DeiT-T/S/B and BERT-B/L)

Constraints

DxPTA Methodology (Section III) Identify the Architecture Parameters of PTA (Section III-A)

tile

Base Architecture & Dataflow of PTA

Our Novel Contributions

accelerator architectures that meet all constraints. It achieves up to 26mm2 area, 4.8W power, 39mJ energy, and 6ms latency across all investigated models, for constraints of 50mm2 area, 5W power, 50mJ energy, and 10ms latency, with 15.2x faster searching time than the exhaustive approach.

core

PE

PE

PE

in the same wavelength λi through the wavelength-division multiplexing (WDM) technique. These signals are passed to the two arms of 50:50 directional coupler (DC) with -90° phase shifter (PS). The outputs of DC (zi0 , zi1 ) are orthogonal in complex plane and can be computed with Eq. 4.      0 1 1 j 1 0 xi zi =√ −jπ/2 j 1 yi zi1 0 e 2 | {z }| {z } (4) DC   PS 1 xi + y i =√ 2 j(xi − yi ) The photodiode (PD) at each output port of DC, converts the signals into photocurrent, which is proportional to the accumulated optical intensities of the input signals. Hence, the output current (I0 ) follows the relation of I0 ∝ ⃗x · ⃗y .

Fig. 3. Our novel contributions in this work.

II. P RELIMINARIES A. Transformer-based Networks A transformer-based network typically consists of multiple identical blocks, known as encoder and decoder blocks. Each block consists of a multi-head self-attention (MHA) module, a feed-forward network (FFN), shortcut connections, as well as a layer normalization (LN) [24]. Furthermore, the decoder block also has cross-attention and masked self-attention modules. The basic encoder block can be formulated as Eq. 1-2. Here, Xl is the input sequences of l-th layer. X̂l+1 = MHA(LN(Xl )) + Xl

(1)

Xl+1 = FFN(LN(X̂l+1 )) + X̂l+1

(2)

Multi-head self-attention (MHA) module has H self-attention heads, where each head transforms the input vector into separate vectors: query (Q), key (K), and value (V) vectors. The attention function between these input vectors can be calculated using Eq. 3. Here, dk is the dimension of Q and K.   QK⊺ Attention(Q, K, V) = sof tmax √ V (3) dk B. Photonic Transformer Accelerator (PTA) In this work, we focus on the LT accelerator [24] as the reference design since it is the state-of-the-art PTA architecture, whose descriptions are provided in the following; see an overview in Fig. 5. The LT Accelerator consists of analog photonic computing elements for accelerating general matrix multiplication (GEMM), optical interconnect for data transmission, and electronics for other operations (e.g., data storage, signal conversion, nonlinear functions, and softmax); see Fig. 5. Following are its key design points. • A single LT chip typically contains Nt tiles, and each tile clusters Nc dynamically-operated photonic tensor cores (DPTCs). • A DPTC is known as the core of LT accelerator, and it contains an array of Nh ×Nv dynamically-operated dotproduct engine (DDot). • A single DDot performs optical dot-product operation between two full-range vectors ⃗x and ⃗y based on coherent interference. Dot-product operation in DDot is performed with the following steps. First, each pair of inputs (xi , yi ) is encoded

III. O UR D X PTA M ETHODOLOGY This methodology identifies the architecture parameters, analyzes the impact of parameters, and develops the constraintaware search algorithm; which are further described below (an overview in Fig. 4). A. Identifying the Architecture Parameters Discussion in Section I-B suggests that the configuration of architecture parameters is important for determining the performance and efficiency of the PTA. Therefore, this step aims to identify parameters that should be customized when designing the PTA. To achieve this, we first analyze the hierarchy of the PTA base architecture and its coherent optical dataflow, and make the following key observations. • Combining the dynamically-operated DDots and the coherent optical dataflow enables multi-wavelength processing which maximizes spectral parallelism and throughput. From such coherent dataflow and operations, we observe some characteristics below. – Multiple wavelengths can be processed in the same DDot unit without requiring prior data programming; see A in Fig. 5. – An Nh ×Nv DDot array is employed to perform multiplications, which defines the parallelism level in a core; see B in Fig. 5. – Multiple cores can be employed to process a chunk of data, which defines the parallelism level in a tile; see C in Fig. 6. – Multiple tiles can be employed to process multiple data chunks, which defines the parallelism level in a chip; see D in Fig. 6. • On-chip global SRAM should have the minimum required size for holding the largest activations in a layer due to layer-by-layer processing, and buffering a portion of offchip data based on the tiling approach (see Fig. 6). Hence, global SRAM size should not be reduced below or increased significantly over this minimum required size, as this will increase the expensive off-chip data access (i.e., high access latency and access energy) or aggravate the static power, respectively [36]–[38].

4

DxPTA Methodology (Section III)

PE

PE

Save the configuration of an architecture with the lowest EDP and meets all constraints

(i.e., DeiT-T/S/B and BERT-B/L)

core

core

core

PE

Nv

core

Determine the significance of the investigated architecture parameters

PTA Architecture with Selected Configuration …

Workloads

Nh

Nt

Identify the prominent parameters of N PTA architecture c

Explore architecture candidates leveraging coherent optical dataflow & parameter significance, then evaluate their costs & EDPs

Photonic Device & Circuit Models

Vary the value of a specific parameter, and then observe its impact

(i.e., area, power, energy, latency)

tile

core

Analyze the characteristics of PTA architecture & its coherent optical dataflow

parameters

PE

core

PE

core

tile

core

Analyze the Impact of Architecture Parameters (Section III-B)

Identify the Architecture Parameters of PTA (Section III-A)

core

significance

Base Architecture & Dataflow of PTA

Constraints

Devise the Constraint-aware Search Algorithm (Section III-C)

core

PE

PE

PE

An artificial general intelligence (AGI)-based application

Fig. 4. Our novel DxPTA methodology with its key steps: (1) identifying of the architecture parameters of PTA; (2) analyzing the impact of architecture parameters; and (3) devising the constraint-aware search algorithm.

M1 SRAM

DPTC

Act. SRAM

M1 SRAM

Act. SRAM

DPTC

Optical Interconnect

Tile 0

Tile Nt-1

Nt

Microcomb

Laser

DDot

DDot

x10 x11 x12 Mod

DDot

DDot

DDot

x20 x21 x22

DDot

DDot

DDot

Mod

B Nh

Nv

DPTC

M1 Dh

×

2

=

𝐼𝐼𝑜𝑜 ∝ 𝑥𝑥⃗ � 𝑦𝑦⃗

× ×

× ×

Tile 0

𝑗𝑗 1

VDD

GND

C

M2

1 𝑗𝑗

1 (𝑥𝑥⃗ + 𝑦𝑦) ⃗ 2

DPTC

Dv

D

× 1

DPTC

Tile 1

× 𝑒𝑒 −𝑗𝑗𝜋𝜋/2

1 𝑗𝑗(𝑥𝑥 ⃗ − 𝑦𝑦) ⃗ 2

DPTC

Tile 0

x2 x x01 𝑥𝑥⃗

−𝑗𝑗𝑦𝑦⃗

Fig. 5. Architecture of the LT accelerator; based on [24]. ×

DDot

Laser

λ2 y2 λ1 y1 λ0 y0 𝑦𝑦⃗

-90o

Microcomb

y02 y12 y22

Mod

λ2 y20

λ0 λ1 λ2 x00 x01 x02 Mod

y01 y11 y21

Mod

λ0 y00

Nλ λ1 y10

Mod

……

DDot

A

Nc

DPTC

DPTC Interconnect ADC Accumulation Unit Out Buffer

Nc Shared M2 Modulation Unit Nc-1 Shared M2 Modulation Unit 0

DPTC

… … … …

DPTC

Shared M2 Modulation Unit

In Buffer DAC Mod.

Shared M2 SRAM

We perform an experimental case study that varies the value of a specific parameter and analyze its impact on different metrics, i.e., area, power, energy, and latency. • Afterward, we evaluate the significance score (S) for each parameter on a specific metric using Eq. 5, whose mechanism is also presented as pseudocode in Alg. 1. Here, si is the ratio between the values of metric-m (i.e., area A or power P ) from architecture with i+1 units and architecture with i units, which represents the impact of unit addition; while K is the total number of ratios. •

Global SRAM

Σ

ADC Accum.

Buffer

Digital Processing Unit

Fig. 6. Data tiling mechanism in the accelerator [24], which partitions data from matrix M1 along the Dv dimension and map them to different tiles.

These observations expose the parameters that should be configured for developing an appropriate accelerator architecture, i.e., number of tiles (Nt ), number of cores-per-tile (Nc ), number of input horizontal waveguides-per-core (Nh ), number of input vertical waveguides-per-core (Nv ), and number of wavelengths (Nλ ). B. Analyzing the Impact of Parameters To ensure the accelerator meets the constraints, an appropriate configuration of parameters (i.e., Nt , Nc , Nh , Nv , and Nλ ) is needed. A promising solution is employing design space exploration (DSE). However, the design space is large due to a high number of possible configurations from different parameter sizes, indicating the need for an optimization. Hence, this step aims to identify the significance of parameters, which will be used to efficiently guide the DSE process. To achieve this, we propose to employ the following steps.

K K 1 X mi+1 units 1 X si = S= K i=1 K i=1 mi units

(5)

with m ∈ {A, P } Experimental results are shown in Fig. 7, from which we make the following key observations. • Increasing parameter size leads to higher power and area, but potentially reduces latency and energy due to higher parallelism. • Nt has the highest impact as each additional tile leads to the highest significance S with 1.26x higher power and 1.24x larger area on average; see E . Introducing a new tile means addition of single/multiple cores and peripherals (e.g., DAC and ADC). • Nc has relatively high impact as each additional core leads to high significance S with 1.23x higher power and 1.20x larger area on average; see F . Introducing a new core means addition of a DDot array with increased size of peripherals (e.g., accumulator). • Nv , Nh , or Nλ have comparable impact to each other, but they are lower than Nt and Nc . For each additional unit of Nv , Nh , or Nλ , power and area are increased by up to 1.16x and 1.06x, respectively; see G . Introducing new DDots within the core or new wavelengths only slightly increases circuit complexity and size (e.g., broadcast unit in the array). C. The Constraint-aware Search Algorithm Based on observations in Section III-B, we develop the following optimization strategy for DSE process. • Exploring Nt and Nc should be performed carefully since their slight changes may incur significant changes on area and power.

5

40 30 20 10 0

Memory Adder PD TIA ADC MZM DAC Laser Area Nt=4; Nc=2; Nh=Nv=12; Nλ varies Nt=4; Nc varies; Nh=Nv=Nλ=12 Nt=4; Nc=2; Nh varies; Nv=Nλ=12 Nt=4; Nc=2; Nv varies; Nh=Nλ=12 80 16 240 16 160 60 80 16 80 G G 60 12 E 180 12 120 45 60 12 60 F G 40 8 120 8 80 30 40 8 40 20 4 60 4 40 15 20 20 4 0 0 0 0 0 0 0 0 0 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10

Nt varies; Nc=2; Nh=Nv=Nλ=12

8 6 4 2 0

Nt varies; Nc=2; Nh=Nv=Nλ=12 lower energy & faster latency

1 2 3 4 5 6 7 8 9 10 # Tiles (Nt)

1.2 0.9 0.6 0.3 0

8 6 4 2 0

Embed

Others Latency Nt=4; Nc=2; Nh=Nv=12; Nλ varies 24 4 4 16 0.8 28 4 lower energy & lower energy & lower energy & lower energy & 18 3 3 0.6 21 12 3 faster latency faster latency faster latency faster latency 2 2 8 0.4 14 2 12 6 1 1 4 0.2 7 1 0 0 0 0 0 0 0 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 # Wavelengths-per-core (Nλ) # Cores-per-tile (Nc) # Horizontal Waveguides-per-core (Nh) # Vertical Waveguides-per-core (Nv)

Nt=4; Nc varies; Nh=Nv=Nλ=12

QKV

Atten

Proj

Nt=4; Nc=2; Nh varies; Nv=Nλ=12

FFN1

FFN2

Head

Nt=4; Nc=2; Nv varies; Nh=Nλ=12

Latency [ms]

Energy [mJ]

(b) Energy Consumption and Latency (for DeiT-B)

Area [mm2]

Power [W]

(a) Power Consumption and Area

Fig. 7. (a) Power and area for different configurations of parameters. (b) Energy consumption and latency for different configurations of parameters when running DeiT-B. We observe similar trends for different workloads (DeiT-T/S and BERT-B/L).

Algorithm 1 Observing the significance of architecture parameters INPUT: Maximum number of observations (J=10); OUTPUT: Significance score of each parameter on area (SA ) and power (SP ); BEGIN Process: 1: for each investigated parameter do 2: Nt = 4; Nc = 2; Nv = 12; Nh = 12; Nλ = 12; // init default values 3: for (j=1; j<(J+1); j++) do 4: set the investigated parameter value with j; 5: cf g[j] = construct(Nt , Nc , Nv , Nh ,Nλ ); 6: A[j], P [j] = eval hw(cf g[j]); 7: i = j-1; 8: if (i>0) then 9: sA [i] = A[i+1]/A[i]; // compute ratio for area 10: sP [i] = P [i+1]/P [i]; // compute ratio for power 11: K = J-1;P K 1 12: SA = K i=1 sA [i]; // compute significance score for area with Eq. P 5 K 1 13: SP = K i=1 sP [i]; // compute significance score for power with Eq. 5 14: return SA , SP ; END

Exploring Nv , Nh , and Nc may be performed more aggressively than Nt and Nc , since their slight changes do not incur significant changes on area and power consumption. • The data dimension is typically evenly sized across layers, which should be leveraged to maximize the resource utilization by selecting the evenly-sized dimension for architecture parameters. We leverage this strategy for devising a constraint-aware search algorithm. Its key ideas are shown in Alg. 2 and discussed below. • We determine a set of values for each parameter, which defines its search space; see Alg. 2: lines 3-10. It is stored in Tcnd , Ccnd , Vcnd , Hcnd , Gcnd for Nt , Nc , Nv , Nh , Nλ , respectively. The search spaces for Tcnd and Ccnd employ incremental values, while the ones for Vcnd , Hcnd and Gcnd are optimized using progressive values with exploration step based on evenly-sized data dimension. • Then, each configuration candidate for the architecture is explored. It is performed by investigating each combination of values from different parameters, and evaluating the area, •

Base Architecture & Dataflow of PTA

Workloads

(i.e., DeiT-T/S/B and BERT-B/L)

DxPTA Methodology Config. Generator PTA HW Simulator Run on GPU

Architecture Configurations

Selector of Arch. Config.

Config. Profiles

EDP calc.

(i.e., area, power, energy, latency)

Constraints

(i.e., area, power, energy, latency)

PTA Architecture w/ Selected Config.

Fig. 8. Experimental setup used in this work.

power, energy, and latency profiles; see Alg. 2: lines 11-14. We evaluate if these profiles meet all constraints and if the energy-delay product (EDP) is lower than the recorded one. If so, then the configuration is saved; see Alg. 2: lines 1522. Here, EDP is the metric for reflecting both performance and energy efficiency. • When the DSE process is finished, the last recorded configuration (cf gsvd ) is selected as the final solution; see Alg. 2: line 23. IV. E VALUATION M ETHODOLOGY •

We implement the DxPTA methodology using PyTorch and run it on the Nvidia RTX 6000 Ada GPU machine; see Fig. 8. For hardware evaluation, we employ the stateof-the-art PTA hardware simulator [24] with 4-bit precision that has been evaluated using the Lumerical Interconnect tools, and incorporate it into the DxPTA. For workloads, we use DeiT-T, DeiT-S, DeiT-B, BERT-B, and BERT-L models. For comparison partners, we consider the LT accelerators (i.e., LT-Base and LT-Large) [24] as the state-of-the-art PTA designs, and an exhaustive approach as the search technique. For constraints, we consider 50mm2 area, 5W power, 50mJ energy, and 10ms latency, to show the applicability of DxPTA for providing solutions under any application requirements. Evaluation metrics include area, power, energy, latency, and search time. V. E XPERIMENTAL R ESULTS AND D ISCUSSION A. Ensuring the PTA Architecture Design to Meet All Design Constraints Fig. 9 provides experimental results of area, power, energy consumption, and latency for different samples of investigated architecture configurations during DSE process across different workloads. These results show that, the state-of-theart accelerators (i.e., LT-Base and LT-Large) do not meet all constraints at once, as they incur significantly larger area

6

Algorithm 2 Our constraint-aware search algorithm INPUT: (1) Maximum number of parameter sizes (Nz =12); (2) Progressive exploration step for non-significant parameters (step=2); (3) Targeted pre-trained network model (net); (4) Constraints for area (constA ), power (constP ), energy (constE ), and latency (constL ); OUTPUT: (1) Final configuration of the architecture parameters (cf gsvd ); BEGIN Initialization: 1: Z1 = []; Z2 = []; 2: EDPsvd = 1000; Process: // Define the search space for each parameter 3: for (nz =1; nz <(Nz +1); nz ++) do 4: Z1 = append(Z1 , nz ); 5: if (nz mod step == 0) then 6: Z2 = append(Z2 , nz ); 7: Tcnd = Z1 ; Ccnd = Z1 ; 8: Vcnd = Z2 ; Hcnd = Z2 ; Gcnd = Z2 ; // Construct and evaluate the configurations 9: for each combination from (∀ nt ∈ Tcnd ), (∀ nc ∈ Ccnd ), (∀ nv ∈ Vcnd ), (∀ nh ∈ Hcnd ), and (∀ nλ ∈ Gcnd ) do 10: cf gcnd = construct(Nt , Nc , Nv , Nh , Nλ ); 11: Acnd , Pcnd = eval hw(cf gcnd ); // evaluate area A and power P 12: Ecnd , Lcnd = eval wload(cf gcnd , net); // evaluate energy E and latency L 13: if (Acnd < constA ) and (Pcnd < constP ) and (Ecnd < constE ) and (Lcnd < constL ) then 14: EDPcnd = calc EDP (Ecnd , Lcnd ); // evaluate EDP 15: if (EDPcnd < EDPsvd ) then 16: EDPsvd = EDPcnd ; 17: cf gsvd = cf gcnd ; 18: return cf gsvd ; // final Nt , Nc , Nv , Nh , and Nλ END

and higher power than the respective constraints; see 1 . Specifically, LT-Base and LT-Large incur about 60mm2 and 112mm2 , respectively (>50mm2 constraint). These sizes are dominated by memory, DAC, and cores; see Fig. 10(a). In terms of power, LT-Base and LT-Large incur about 15W and 28W power, respectively (>5W constraint). These power values are dominated by Mach-Zender Modulator (MZM), DAC, photodetector, and ADC; see Fig. 10(b). The reason is that, LT-Base and LT-Large designs have fixed configurations for accelerating diverse workloads, thus they may not be applicable for different requirements. In contrast, our DxPTA consistently finds the suitable configurations that meet all constraints and have the lowest EDP scores among the candidates across different workloads; see 2 for power and area, 3 for energy consumption, 4 for latency, and 5 for EDP. The reason is that, DxPTA employs a search algorithm that incorporates all constraints in its exploration process, ensuring the selected configuration to fulfills the requirements. Consequently, the DxPTAgenerated accelerators significantly reduce the area and power as compared to LT-Base and LT-Large, by achieving up to 76.9% area saving and 82.7% power saving; see 6 in Fig. 10. Furthermore, our DxPTA-generated accelerators also achieve comparable area and power consumption to the accelerators

whose configurations generated from exhaustive search (i.e., Exh-DeiT and Exh-BERT); see 7 in Fig. 10. The reason is that, DxPTA already considers the significance of parameters in its search strategy to ensure the coverage of potential configuration candidates in the DSE process. Hence, DxPTA can find configurations that are close to the ones from the exhaustive search. B. Enabling High-Performance and Energy-Efficient Transformer Inference Fig. 11 presents the performance (i.e., FPS: frame-persecond) and energy consumption of DeiT-B processing using different platforms: CPU, GPU, electronic accelerators, state-of-the-art PTAs, and our DxPTA-based PTA. These results show that, the DxPTA-based PTA achieves comparable performance (FPS) to the LT-Base and LT-Large designs, and provides significant improvements from conventional platforms; see 8 . Specifically, the DxPTA-based PTA improves the performance by 189x from CPU, 4.1x from GPU, 20.1x from AutoVit-4, and 17.2x from HeatVIT-8. Such remarkable performance improvements come from the high-speed nature of light propagation. The DxPTA-based PTA also achieves comparable energy-efficiency to the LTBase and LT-Large designs, and provides significant energy savings from conventional platforms; see 9 . Specifically, the DxPTA-based PTA saves energy consumption by 782.1x from CPU, 15.2x from GPU, 31.6x from AutoVit-4, and 27.6x from HeatVIT-8. Such energy efficiency improvements come from its lightweight optical-based processing and minimum energy consumption from electronic parts. Note, all these improvements are achieved by DxPTA while meeting all given constraints at once, further highlighting the benefits of our DxPTA methodology. C. Architecture Searching Time Speedup Fig. 12 presents the searching time of the exhaustive approach and the guided search in our DxPTA across different workloads. These results highlight that DxPTA achieves a significant searching time speedup, i.e., by 15.2x faster than the exhaustive one. This speedup comes from the optimized search space in DxPTA through exploration steps, guided by data dimension, coherent optical dataflow, and parameter significance. This speedup is beneficial to make the DxPTA methodology an efficient and scalable solution for developing appropriate PTA architectures under different possible design constraints. VI. C ONCLUSION We propose a novel DxPTA methodology to perform DSE for enabling efficient HW/SW co-design of the PTA architecture that meets all given constraints. DxPTA identifies the prominent architecture parameters based on the coherent optical dataflow, analyzes the significance of parameters, and then devises a constraint-aware search algorithm. Experiments show that, DxPTA successfully finds the appropriate PTA architectures for different transformer models, achieving up

7

Configuration that does not meet all constraints

Configuration that meets all constraints

Configuration found by DxPTA that meets all constraints and has the lowest EDP

(2) DeiT-S

(1) DeiT-T

(5) BERT-L (3) DeiT-B (4) BERT-B 80 80 80 80 80 constA constA constA constA constA (a.1) (a.2) (a.3) (a.4) (a.5) 60 60 60 60 60 1 LT-Large 1 LT-Large 1 LT-Large 1 LT-Large 1 LT-Large LT-Base 40 LT-Base LT-Base LT-Base 40 40 LT-Base 40 40 2 2 2 2 2 20 constP 20 20 constP constP 20 constP 20 constP 0 0 0 0 0 100 150 200 250 8 0 (b) 2.0 0 50const 50const 100 150 200 250 32 0 50const 100 150 200 250 20 0 50const 100 150 200 250 200 0 50const 100 150 200 250 (b.1) (b.2) A A (b.5) A A (b.4) A (b.3) 6 15 150 24 1.5 100 10 16 4 1.0 constE 50 5 8 0.5 2 3 3 0 0 0 0.0 0 3 3 3 50const 100 150 200 250 4 0 50const 100 150 200 250 20 0 50const 100 150 200 250 3.2 0 100 150 200(c.1) 250 1.2 0 (c) 0.4 0 50const 50const 100 150 200(c.2) 250 A (c.5) A A (c.4) A (c.3) A 15 2.4 3 0.9 0.3 constL 4 10 1.6 4 0.2 2 0.6 4 4 5 0.8 0.1 1 0.3 4 0 0 0 0 0 100 0 6.0 0 0 50const 100 150 200 250 3200 40 0 50 100 150 200 250 100 150 200 250 (d) 0.4 0 50const 50 100 150 200 250 50 100 150 200 250 const const const (d.1) A A A (d.4) A (d.3) (d.5) (d.2) A 0.3 75 4.5 30 2400 0.2 3.0 50 20 1600 5 5 5 5 5 0.1 1.5 25 10 800 0.0 0 0 0 0 0 50 100 150 200 250 0 50 100 150 200 250 0 50 100 150 200 250 0 50 100 150 200 250 0 50 100 150 200 250 Area [mm2] Area [mm2] Area [mm2] Area [mm2] Area [mm2]

Thousands

Thousands

Thousands

EDP [x10-6 Js]

Latency [ms]

Energy [mJ]

Power

Thousands

Thousands [ W]

(a)

Area [mm2]

125 100 75 50 25 (a) 0 30 25 20 15 10 5 0

constA

area savings 6

Power [W] Thousands

7

constP

power savings 6 7

(b)

Micro_comb Adder TIA MZM Laser

Buffer Memory Core ADC DAC

Buffer Memory

Adder

Photodetector

TIA

ADC

MZM

DAC

Laser

(a)

100 10 1

FPS improvements

8

10000

Energy [mJ] (log scale)

Performance [FPS] (log scale)

1000

1000 100 10

100 10 1 0.1 0.01

DxPTA expedites the architecture searching time by 15.2x faster than the exhaustive approach

faster

DeiT-T

faster

DeiT-S

faster

DeiT-B

Exhaustive faster

BERT-B

DxPTA faster

BERT-L

Fig. 12. Searching time profiles of the exhaustive search approach and the searching strategy in our DxPTA.

Fig. 10. Experimental results of (a) area and (b) power for LT-Base, LTLarge, exhaustive search-based accelerators for DeiT/BERT models (i.e., ExhDeiT/BERT), and DxPTA-based accelerators for DeiT/BERT models (i.e., DxPTA-DeiT/BERT). 10000

Search Time Normalized to Exhaustive DeiT-T (log scale)

Fig. 9. Experimental results of (a) power vs. area, (b) energy vs. area, (c) latency vs. area, and (d) EDP vs. area for different samples of configurations during DSE process, across different workloads: DeiT-T, DeiT-S, DeiT-B, BERT-B, and BERT-L.

energy savings 9

1

(b)

Fig. 11. Experimental results of (a) performance and (b) energy consumption of DeiT-B processing using CPU, GPU, electronic accelerators (i.e., AutoViT4b [25] and HeatViT-8b [26]), state-of-the-art PTAs (i.e., LT-Base and LTLarge [24]), and DxPTA.

to 26mm2 area, 4.8W power, 39mJ energy, and 6ms latency, for constraints of 50mm2 area, 5W power, 50mJ energy, and 10ms latency; with 15.2x faster search time than the exhaustive approach. These results demonstrate the potential of DxPTA methodology for enabling efficient PTA design automation for diverse AGI-based applications. R EFERENCES [1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez et al., “Attention is all you need,” Advances in Neural Information Processing Systems (NIPS), vol. 30, no. 1, pp. 261–272, 2017.

[2] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021. [3] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML). PMLR, 2021, pp. 10 347–10 357. [4] S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM Computing Surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022. [5] K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao, “A survey on vision transformer,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 45, no. 1, pp. 87–110, 2023. [6] A. Mumuni and F. Mumuni, “Large language models for artificial general intelligence (agi): A survey of foundational principles and approaches,” arXiv preprint arXiv:2501.03151, 2025. [7] G. Yenduri, R. Murugan, P. Kumar Reddy Maddikunta, S. Bhattacharya, D. Sudheer, and B. Bhushan Savarala, “Artificial general intelligence: Advancements, challenges, and future directions in agi research,” IEEE Access, vol. 13, pp. 134 325–134 356, 2025. [8] S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the transformer,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC). IEEE, 2020, pp. 84–89. [9] P. Qi, E. H.-M. Sha, Q. Zhuge, H. Peng, S. Huang, Z. Kong, Y. Song, and B. Li, “Accelerating framework of transformer by hardware design and model compression co-optimization,” in 2021 IEEE/ACM International Conference On Computer Aided Design (ICCAD), 2021, pp. 1–9. [10] H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021, pp. 97–110. [11] M. Sun, H. Ma, G. Kang, Y. Jiang, T. Chen, X. Ma, Z. Wang, and Y. Wang, “Vaqf: Fully automatic software-hardware co-design framework for low-bit vision transformer,” arXiv preprint arXiv:2201.06618, 2022.

8

[12] M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memorybased acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, pp. 1071–1085. [13] N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Young, X. Zhou, Z. Zhou, and D. A. Patterson, “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA), 2023. [14] H. You, Z. Sun, H. Shi, Z. Yu, Y. Zhao, Y. Zhang, C. Li, B. Li, and Y. Lin, “Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2023, pp. 273–286. [15] F. P. Sunny, E. Taheri, M. Nikdast, and S. Pasricha, “A survey on silicon photonics for deep learning,” ACM Journal of Emerging Technologies in Computing System (JETC), vol. 17, no. 4, pp. 1–57, 2021. [16] K. Shiflett, A. Karanth, R. Bunescu, and A. Louri, “Albireo: Energyefficient acceleration of convolutional neural networks via silicon photonics,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2021, pp. 860–873. [17] B. J. Shastri, A. N. Tait, T. Ferreira de Lima, W. H. Pernice, H. Bhaskaran, C. D. Wright, and P. R. Prucnal, “Photonics for artificial intelligence and neuromorphic computing,” Nature Photonics, vol. 15, no. 2, pp. 102–114, 2021. [18] J. Gu, C. Feng, H. Zhu, R. T. Chen, and D. Z. Pan, “Light in ai: toward efficient neurocomputing with optical neural networks—a tutorial,” IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 69, no. 6, pp. 2581–2585, 2022. [19] Z. Yin, M. Zhang, N. Gangi, R. Huang, J. Zhang, and J. Gu, “Simphony: A device-circuit-architecture cross-layer modeling and simulation framework for heterogeneous electronic-photonic ai system,” in 2025 62nd ACM/IEEE Design Automation Conference (DAC). IEEE, 2025, pp. 1–7. [20] Y. Shen, N. C. Harris, S. Skirlo, M. Prabhu, T. Baehr-Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englund et al., “Deep learning with coherent nanophotonic circuits,” Nature photonics, vol. 11, no. 7, pp. 441–446, 2017. [21] A. N. Tait, T. F. De Lima, E. Zhou, A. X. Wu, M. A. Nahmias, B. J. Shastri, and P. R. Prucnal, “Neuromorphic photonic networks using silicon photonic weight banks,” Scientific Reports, vol. 7, no. 1, p. 7430, 2017. [22] F. Sunny, A. Mirza, M. Nikdast, and S. Pasricha, “Crosslight: A crosslayer optimized silicon photonic neural network accelerator,” in 2021 58th ACM/IEEE design automation conference (DAC). IEEE, 2021, pp. 1069–1074. [23] J. Feldmann, N. Youngblood, M. Karpov, H. Gehring, X. Li, M. Stappers, M. Le Gallo, X. Fu, A. Lukashchuk, A. S. Raja et al., “Parallel convolutional processing using an integrated photonic tensor core,” Nature, vol. 589, no. 7840, pp. 52–58, 2021. [24] H. Zhu, J. Gu, H. Wang, Z. Jiang, Z. Zhang, R. Tang, C. Feng, S. Han, R. T. Chen, and D. Z. Pan, “Lightening-transformer: A dynamicallyoperated optically-interconnected photonic transformer accelerator,” in 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2024, pp. 686–703. [25] Z. Li, M. Sun, A. Lu, H. Ma, G. Yuan, Y. Xie, H. Tang, Y. Li, M. Leeser, Z. Wang et al., “Auto-vit-acc: An fpga-aware automatic acceleration framework for vision transformer with mixed-scheme quantization,” in 2022 32nd International Conference on Field-Programmable Logic and Applications (FPL). IEEE, 2022, pp. 109–116. [26] P. Dong, M. Sun, A. Lu, Y. Xie, K. Liu, Z. Kong, X. Meng, Z. Li, X. Lin, Z. Fang, and Y. Wang, “Heatvit: Hardware-efficient adaptive token pruning for vision transformers,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2023, pp. 442–455. [27] Y. Li, A. Louri, and A. Karanth, “Sprint: A high-performance, energyefficient, and scalable chiplet-based accelerator with photonic interconnects for cnn inference,” IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 10, pp. 2332–2345, 2022. [28] ——, “Spacx: Silicon photonics-based scalable chiplet accelerator for dnn inference,” in 2022 IEEE International Symposium on HighPerformance Computer Architecture (HPCA), 2022, pp. 831–845. [29] S. Afifi, F. Sunny, M. Nikdast, and S. Pasricha, “Tron: Transformer neural network acceleration with non-coherent silicon photonics,” in Great Lakes Symposium on VLSI (GSVLSI) 2023, 2023, pp. 15–21.

[30] S. Afifi, O. Alo, I. Thakkar, and S. Pasricha, “A light-speed large language model accelerator with optical stochastic computing,” in Great Lakes Symposium on VLSI (GLSVLSI) 2025, 2025, pp. 922–928. [31] ——, “Astra: A stochastic transformer neural network accelerator with silicon photonics,” ACM Transactions on Embedded Computing Systems (TECS), 2025. [32] Y. Li, A. Louri, and A. Karanth, “Merit: A sustainable dnn accelerator design with photonic phase-change memory,” IEEE Transactions on Sustainable Computing (TSUSC), vol. 10, no. 4, pp. 705–716, 2025. [33] H. Li, D. Chen, and T. Mitra, “Hyatten: Hybrid photonic-digital architecture for accelerating attention mechanism,” in 2025 Design, Automation & Test in Europe Conference (DATE), 2025, pp. 1–7. [34] W.-T. Chang, C.-F. Wu, and Y.-C. Lo, “P-dac: Power-efficient photonic accelerators for llm inference,” in 2025 62nd ACM/IEEE Design Automation Conference (DAC), 2025, pp. 1–7. [35] H. Zhu, Z. Zhou, S. Ning, X. Wu, R. Chen, Y. Wan, and D. Pan, “Enlighten: Lighten the transformer, enable efficient optical acceleration,” arXiv preprint arXiv:2510.01673, 2025. [36] R. V. W. Putra, M. A. Hanif, and M. Shafique, “Drmap: A generic dram data mapping policy for energy-efficient processing of convolutional neural networks,” in 2020 57th ACM/IEEE Design Automation Conference (DAC), 2020, pp. 1–6. [37] ——, “Romanet: Fine-grained reuse-driven off-chip memory access management and data organization for deep neural network accelerators,” IEEE Transactions on Very Large Scale Integration Systems (TVLSI), vol. 29, no. 4, pp. 702–715, 2021. [38] ——, “Pendram: Enabling high-performance and energy-efficient processing of deep neural networks through a generalized dram data mapping policy,” arXiv preprint arXiv:2408.02412, 2024.

Record · ID 266151 · SHA-256 1fed785b7760a12d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.