ConceptioArchivearXiv CS
arXiv CSopen access

Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained Optimization

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

1

Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained Optimization Xiangming Huang, Guannan Zhang, Lu Lu, Raphaël Pestourie

This is the accepted manuscript version for publication by IEEE. © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

arXiv:2607.07682v1 [cs.LG] 8 Jul 2026

I. A BSTRACT The inverse design of physical systems governed by partial differential equations is computationally demanding due to the high dimensionality and non-convexity of design spaces. Generative models for inverse design often lack robustness and transferability, whereas evolutionary strategies are robust but struggle in high-dimensional spaces. This paper introduces a Neural Operator–enabled Topology-informed Evolutionary Strategy (NOTES) that integrates dimensionality reduction, representation learning, and evolutionary optimization for efficient and transferable inverse design. NOTES couples a DeepONet-based neural operator with the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) to perform global optimization in a compact latent space that encodes topologyaware priors while discovering high-performance designs for unseen operating conditions. Applied to nanophotonic beamdeflector inverse design governed by Maxwell’s equations, NOTES reduces the design dimensionality from 256 to 25 and consistently achieves over 95% efficiency, outperforming CMA-ES, topology optimization, and other baselines. Applied to structural optimization, NOTES discovers designs that achieve compliance down to 246. By decoupling topology learning of a DeepONet from the governing physics in a PDE solver, NOTES provides a flexible and transferable framework for the inverse design of physical systems. II. I NTRODUCTION We introduce a Neural Operator-enabled Topologyinformed Evolutionary Strategy (NOTES) that integrates implicit neural reparameterization and evolutionary optimization to provide an efficient, flexible, and transferable framework for inverse design of physical systems. The method unifies a DeepONet-based neural operator [1] with the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) [2] to perform global optimization in a compact latent space, which encodes topology-aware priors and reduces the number of optimization variables. Although NOTES does not embed the governing partial-differential equation (PDE) explicitly into the neural operator loss, the framework remains physicsinformed through the construction of its training dataset. The DeepONet is trained exclusively on designs obtained from

direct PDE-constrained optimization of the target physical systems, so the learned latent representation is biased toward geometries that encode physically meaningful high-performance features. In this sense, physical information enters NOTES indirectly through the data distribution rather than through a residual-based loss function [3]. In addition, DeepONet allows architectural inductive biases to be incorporated into the representation, illustrated here through binarization constraints and, in prior work, through hard enforcement of physical priors such as boundary conditions and invariance [4]. NOTES is particularly well suited for PDE-constrained inverse design problems where relevant high-performance designs are already available, allowing transferable geometric features to be learned from existing data. By combining representation learning with evolutionary optimization, the framework aims to make evolutionary strategies practical for high-dimensional topology optimization problems that are otherwise prohibitively expensive for direct search. In addition to improving optimization efficiency, NOTES provides a robust mechanism for consistently generating diverse highperformance designs across different physical settings. III. P RIOR W ORK A. Evolutionary Strategy for Inverse Design CMA-ES [2], [5]–[9] is a widely used derivative-free evolutionary strategy [10] for nonconvex optimization because of its robustness and flexibility. However, its performance deteriorates rapidly in high-dimensional search spaces due to the curse of dimensionality [11]. Recent works [12]–[14] therefore explore dimensionality reduction and latent-space optimization to improve scalability. B. Representation Learning for Inverse Design Representation learning has been widely used to accelerate PDE-constrained inverse design [15], [16]. Several prior works [17]–[19] constructed various generative models for inverse design. Variational Autoencoders (VAEs) are widely used for this purpose across many applications [18], [20]– [24]. However, controlling inductive bias in VAEs remains challenging because it can conflict with the inherent bias introduced by the variational formulation [25], [26]. Our method follows the line of work that couples neural representations with evolutionary optimization [13], [27]. C. Neural Operator Although DeepONet has been widely used in topology optimization and PDE learning [1], [28]–[33], it is typically

2

Operator-enabled Topology-informed A Neural Evolutionary Strategy (NOTES)

A

Nanophotonic

Topology

Topology Neural Operator

Before

B Optimization Topology

Latent space

PDE Solver

C

dim = 25 dim = 256

illconditioned

Si

Inductive bias

SiO2

Material choice

λ

Coordinate

Trunk Net

meshless

employed as a surrogate PDE solver that maps PDE coefficients to solution fields. In contrast, our framework uses DeepONet as a nonlinear topology representation that maps a low-dimensional latent space to the design topology. The latent space captures transferable geometric features from highperformance structures, enabling compact and reusable design representations. Our approach is further inspired by prior studies on binarization bias in neural networks [34], implicit neural reparameterization [35], [36], transferability of neural operators [32], and periodicity enforcement in DeepONet [37]. The remainder of this manuscript is organized as follows. Section IV presents the overall NOTES framework and optimization procedure. Section V introduces two PDEconstrained inverse design applications used to evaluate the proposed method. Section VI presents comparisons with multiple baselines and ablation studies, including evaluations of CMA-ES as the optimizer, DeepONet as the nonlinear decoder, transferability across PDE settings, robustness to training randomness, and the importance of physics-guided training data. IV. M ETHOD A. Schematic Overview NOTES addresses topology optimization with PDE constraints [15]. Such optimization problems are often formulated as minimization problems. When a higher value of a figure of merit (FOM) is desired, the objective is reformulated as minimizing the negative of the FOM: s.t.

F(x; v) = 0

900

Photonics Mechanics

v∈C

0

C PDE equations:

Fig. 1. (A) Overview of the neural operator-enabled topology-informed evolutionary algorithm. Using a latent space as the search space, topologies are generated by a neural operator. The target property of the generated topology is evaluated by a PDE solver. The CMA-ES algorithm is applied in the latent space to maximize the property. (B) Before the optimization, PCA reduces the high-dimensional design space to a compact latent space. The latent space captures key geometric features and provides a compact representation of the design space. (C) A neural operator, based on DeepONet [1] is trained to map latent vectors to their corresponding design topologies. The architecture and loss function of the neural operator encourage the generated design to satisfy binarization constraints.

min −FOM(x; v)

1100

75

1000

Binarization

topological insights

Dimensionality Reduction 25 principal components

wavelength

Branch Net

PCA

D

Scalar property

Topology Neural Operator

DeepONet Latent vector

Full-view MBB Beam

θ

Covariance matrix adaptation evolution strategy Latent space

B

(1)

-100

-50

0

100

800

Fig. 2. Illustration of the applications and dimensionality reduction framework used in NOTES. (A) Side view of the 1D nanophotonic beam deflector. An incident TM-polarized plane wave with wavelength λ propagates from the bottom, and the normalized transmitted power in the θ direction is computed as the deflection efficiency. (B) Full-view of the 2D Messerschmitt–Bölkow– Blohm (MBB) beam for structural optimization. A downward force is applied at the center, while both lower corners are supported. Using symmetry, only half of the domain is optimized. The objective is to minimize the elastic compliance (elastic potential energy) of the structure. (C) In both applications, the target property is a scalar quantity evaluated by solving the corresponding PDEs shown in the figure. (D) PCA is used for dimensionality reduction. We illustrate with the projection of 6129 high-performance photonics designs onto the first two principal components that reveals clustering behavior with respect to the incident wavelength, indicating that PCA captures meaningful geometric features of the design manifold.

where FOM is a scalar objective function, F denotes the PDE operator, and C represents the feasible space. Additional constraints may be added to enforce manufacturability [38]. x is the PDE solution, and v is the design variable. Solving for the optimal design v∗ is challenging because the objective is typically nonconvex and defined over a high-dimensional design space. In addition, evaluating the objective requires repeatedly solving the underlying PDE system, which is computationally expensive. NOTES consists of two stages: training and optimization. During the training stage, we construct a latent representation from high-performance designs using Principal Component Analysis (PCA) and train a neural operator to map latent vectors to design topologies. During the optimization stage, illustrated in Figure 1A, the trained neural operator acts as a geometric reparameterization of the design space, while CMAES performs optimization over the learned latent representation.

B. Dimensionality Reduction We construct a low-dimensional latent space from the original design space using dimensionality reduction and train a neural network decoder that maps latent vectors to design topologies: v = N N (u),

dim(u) = n ≪ dim(v) = N

(2)

3

As illustrated in Figure 1B, the latent space captures key topological features while significantly reducing the dimensionality of the optimization problem. The PDE-constrained optimization problem is therefore reformulated as min

− FOM(x; N N (u))

s.t.

F(x; N N (u)) = 0,

u∈Rn

(3)

N

N N (u) ∈ R . Optimizing in the latent space mitigates the curse of dimensionality [11] and improves the scalability of evolutionary optimization [39]. C. Representation Learning We train a DeepONet-based neural operator that maps latent vectors and density-function coordinates to design topologies (Figure 1C). Using this neural representation, we reformulate the optimization problem in the latent space: min

− FOM(x; v̂)

s.t.

F(x; v̂) = 0,

u∈Rn

v̂ ∈ RN ,   v̂ = σ G(u, y); η̂, β̂ .

(4)

Here, G(u, y) denotes the DeepONet used to output the design vector v̂. Its architecture consists of two sub-networks: the branch net g(.), which encodes the latent vector u = (u1 , . . . , un ), and the trunk net f (.), which encodes the density-function coordinates y in the design. To facilitate training, y is normalized between 0 and 1. Formally, DeepONet can be expressed as: (5) G(u, y) = g(u1 , . . . , un ) ⊙ f (y) , {z } |{z} | branch trunk where ⊙ denotes the component-by-component (Hadamard) product. Design topologies are parameterized as binary sequences because intermediate material states are nonphysical. To encourage binarized outputs, we apply a sigmoid transformation at the output layer of DeepONet [34]: σ(û; β, η) =

1 1 + exp(−β û − η)

(6)

where β controls the degree of binarization from a threshold η. The model is trained using binary cross-entropy loss together with a binarization regularization term:

LBCE = −

N  1 X T v log v̂ + (1 − v)T log(1 − v̂) N i=1

Lbias = (1N − v̂)T v̂ L = LBCE + αLbias

(7) (8) (9)

Here, α controls the contribution of the binarization regularization term and 1N denotes a N-dimensional vector of ones. Although binarization is enforced as a soft constraint,

progressively increasing β during training yields increasingly binarized topologies. Note that fabrication-aware priors are incorporated implicitly through the neural representation and the distribution of high-performance training designs rather than through the final optimization objective. D. Optimization Over Latent Space Given a trained neural operator, CMA-ES is applied to solve the optimization problem defined in Eq. 4. The latent vector u is initialized through random sampling in the latent space, with the sampling scale determined from the latent-coordinate distribution of the training data. During optimization, each generated topology is evaluated using a PDE solver to compute the scalar objective function. The PDE solver ensures accurate physics evaluation and is the dominant computational cost in NOTES. V. A PPLICATION A. Problem Setup To demonstrate both the effectiveness and applicability of NOTES across physical domains, we evaluate the proposed framework on two PDE-constrained inverse design applications with fundamentally different governing physics and geometric structures. The first is a one-dimensional nanophotonic inverse design, where the objective is to generate high-efficiency metagrating patterns for prescribed incident wavelengths and deflection angles. The second is a twodimensional structural inverse design, where the objective is to generate low-compliance load-bearing structures under prescribed design resolutions and material volume constraints. The nanophotonic application serves as the primary benchmark throughout the main manuscript because it provides a well-established inverse design setting with strong competing baselines, while the structural mechanics application is used to further verify that NOTES is transferable across different PDEs or physical domains and effective in a more complex design space. We perform the inverse design of a beam deflector, following settings similar to Ref. [40]. The objective is to maximize the deflection efficiency (the ratio of the electromagnetic power deflected to the incident light power) of the +1 diffraction order in silicon nanoridge metagratings. The grating thickness is fixed at 325 nm, and the incident light is TMpolarized. The grating period equals the incident wavelength divided by the sine of the deflection angle in magnitude and is discretized into 256 equal ridges, each assigned one of two material states. Figure 2 A illustrates a single grating period. Each period is parameterized as a 256-dimensional binary vector, where 1 denotes silicon and 0 denotes air. A differentiable PDE solver, Meent [41], is then used to solve the PDE in Figure 2 C and to evaluate the objective, with the incident wavelength λ and deflection angle θ as parameters: min

v∈{0,1}256

−P (v; λ, θ)

(10)

The deflection efficiency is a percentage, denoted as P and dependent on E (Fig. 2C). The permittivity profile, denoted

4

as v, is a binary vector. The objective is to find the optimal v∗ that maximizes P . As in Eq. 4, we use the optimization formulation with our DeepONet representation on a latent space of dimension n = 25. The second application considers a two-dimensional structural topology optimization problem. We apply NOTES to the classical Messerschmitt–Bölkow–Blohm (MBB) beam benchmark on a design domain of size 192 × 64, resulting in 12,288 pixel-level design variables, illustrated in Fig. 2B. The overall problem setup and finite-element implementation follow the framework in Ref. [35]. The discrete compliance minimization problem is given by min

U T KU

subject to

KU = F,

v∈R192×64

(11)

V (v) = V0 , 0 ≤ vij ≤ 1,

∀ (i, j),

where vij denotes the physical density at each pixel, K(v) is the global stiffness matrix, F is the external load vector, and U (K, F ) is the corresponding displacement vector obtained from the finite-element equilibrium equation. The total material volume V (v) is constrained to a predefined volume fraction V0 . The following constrained sigmoid transformation is applied after the forward pass in the DeepONet to enforce the volume and density constraints in Equation 11: 1 , 1 + exp[−v̂ij − b(v̂, V0 )] V (ṽ) = V0 ,

ṽij = such that

(12)

where b(v̂, V0 ) is a scalar bias term determined by binary search so that the transformed output field satisfies the predefined volume fraction. Finally, the physical density field v is obtained by applying a cone filter with a radius of 2 to ṽ, which smooths isolated small features as in Ref. [35]. As in Eq. 4, we use the optimization formulation with our DeepONet representation on a latent space of dimension n = 60. B. Training Data and Latent Space For the nanophotonic beam deflector application, we use the publicly available MetaNet benchmark dataset [42], which contains high-performance grating designs generated by GLOnet and adjoint-based topology optimization [40] under multiple (wavelength, deflection angle) operating conditions. From this dataset, we retain only designs whose deflection efficiency exceeds 0.9 for their corresponding operating pair (λ, θ). A practical difficulty is that the deflection efficiency is not uniformly distributed across all operating conditions. Several wavelength–angle pairs contain no high-efficiency designs. For instance, under (λ, θ) = (1100 nm, 40◦ ), no design with efficiency above 0.9 is available in the database, and the best designs selected from other operating conditions achieve only about 0.75 efficiency. Therefore, the neural operator cannot rely solely on direct memorization of condition-specific

optimal patterns. Instead, it must learn transferable topological features from high-efficiency designs to generate new candidate structures for previously underrepresented operating conditions. To construct a latent space for DeepONet in photonics, we applied PCA [43] to the dataset of 6129 high-performance designs spanning different wavelengths and deflection angles, following the procedure outlined in Metanet [42] that conserves the translational invariance of the periodic grating. The physical parameters (λ, θ) are not directly encoded in the latent space, only indirectly influencing the latent representation through their corresponding high-performance geometries in the dataset. The distribution of data points along the first two principal components is visualized in Figure 2D, which reveals distinct clustering by wavelength. The visualization indicates that the latent space captures distinct geometric features associated with different wavelengths and deflection angles. The first 25 principal components explain more than 90% of the total variation and are chosen as our lowdimensional latent space. With this latent space, DeepONet learns a 25-dimensional representation that generates 256dimensional designs. For the MBB beam application, the dataset is generated using density-based topology optimization with the L-BFGS optimizer. The designs are generated with a prescribed volume fraction of 0.4. To improve diversity in the generated dataset, each new optimization is initialized from the previous design after applying erosion operations with varying rates. The resulting dataset has a median compliance of approximately 270 (SI section 2.1 & 2.2). We retain 857 optimized designs as base samples and further augment the dataset using additional erosion and dilation operations (Figure S8 in SI). For each base design, two augmented variants are generated. After augmentation, the complete dataset contains 2571 designs, which are randomly split into 2056 training samples and 515 testing samples. To obtain the latent space for the MBB beam application, we fit PCA on the training samples and retain the first 60 principal components, which explain approximately 83% of the variance in the training dataset (Figure S9 in SI). C. Training and Optimization The DeepONet implementation is provided by DeepXDE [44]. Using a 60/40 train/test split, the model is trained on 3677 samples for the photonic case. Both the branch and trunk networks contain three hidden layers with 60 neurons per layer, resulting in approximately 20k trainable parameters. A sigmoid activation is applied at the output with an adaptive scale factor and a threshold initialized to the average value of the training data. The network is trained using the ADAM optimizer with learning rate 10−3 for 400,000 epochs. The final validation MSE is approximately 0.06 (Figure S2 in SI). The trained DeepONet achieves high reconstruction quality (Figure S3 in SI). The PDE solver is provided by Meent [41], which implements differentiable rigorous coupled-wave analysis (RCWA). This enables both gradient-based optimization and evolutionary optimization in the learned latent space. CMA-ES is implemented using pycma [45].

5

A

Efficiency: 0.95

B

Efficiency: 0.94

B

A NOTES Frequency

7

C

Efficiency: 0.94

D

efficiency: 95%

4

0

Efficiency: 0.94

0.90 0.86 efficiency

0.82

C

0.94

Efficiency: 0.94

Frequency

F

Aspect Ratio: 65.52

D

CMA-ES alone 6

E

Best design from NOTES

Best design from CMA-ES alone efficiency: 94%

4

Efficiency: 0.94 0

0.4

E

0.5

efficiency

0.8

Aspect Ratio: 65.52

0.9

F Best design from L-BFGS + DeepONet efficiency: 95%

L-BFGS + DeepONet

Fig. 3. (A-F) qualitatively present six optimized metagrating patterns generated by NOTES for the 1100 nm wavelength and 60° deflection angle. Although the best design in the training data achieves an efficiency equal to 0.9 for this setting, the method consistently produces results with higher efficiency. This shows the ability of NOTES to discover designs beyond the training data. Most generated designs satisfy binarization constraints. (E, F) show designs with minor imperfections. These cases are rare (3%) across all operating conditions. The aspect ratios of the six designs are 65.52, 16.38, 16.38, 16.38, 16.38, and 21.84, respectively. High-aspect-ratio structures are uncommon, with 95% of the generated designs having aspect ratios below 33.

Preliminary optimization results suggest that including a binarization term in the loss function is helpful. Approximately 20% of the generated designs remain partially non-binary without binarization bias, while incorporating a binarization bias reduces this fraction to about 3%. The latent vector is initialized randomly within the PCA coordinate range observed in the training dataset. CMA-ES is applied with a population size of 20 and a maximum of 100 iterations. Different initial sampling radii are used for the baseline CMA-ES and for NOTES (σ0 = 10 and σ0 = 1000, respectively), resulting from hyperparameter tuning. The initial standard deviation for NOTES is relatively large because the PCA latent coordinates are not normalized. For the MBB beam application, the number of parameters in the DeepONet is increased to 1.6 million because of the output complexity (SI section 2.3). The training procedure follows the same setup as the 1D photonic case, including the use of a custom sigmoid activation and a binarization bias in the loss function. However, unlike the photonic application, the optimizer enforces the volume constraint strictly instead of the DeepONet. The trained DeepONet achieves high reconstruction quality for the 2D mechanical case (Figure S10 in SI). VI. R ESULTS A. Baselines Configuration For the nanophotonic application, we compare NOTES against four baselines: topology optimization, GLOnet [40],

Frequency

7

4

0

0.55

0.65

efficiency

0.95

Aspect Ratio: 17.00

Fig. 4. This figure compares NOTES with the standard CMA-ES and LBFGS + DeepONet for the 1100 nm wavelength and 60° deflection-angle setting. Comparing panels (A), (C), and (E), NOTES exhibits substantially lower variance than the other two methods, with most generated designs achieving higher than 92% efficiency. The aspect ratios of designs (B), (D), and (F) are 65.52, 65.52, and 17.00, respectively, with panel (F) containing intermediate nonphysical material values.

CMA-ES, and L-BFGS + DeepONet. Topology optimization performs direct gradient-based optimization of the highdimensional design. GLOnet [40] is a sampling-based inversedesign framework that previously achieved state-of-the-art performance for this application. Standard CMA-ES performs direct evolutionary optimization in the 256-dimensional topology. Finally, L-BFGS + DeepONet applies the L-BFGS optimizer in the same latent space and neural representation as NOTES, differing only in the choice of optimizer. For the mechanical application, the baselines are L-BFGS-based topology optimization in high-dimensional space and the LBFGS + DeepONet baseline. Due to the 12,288-dimensional feasible space, direct CMA-ES could not achieve comparable performance. In both applications, we also compare NOTES to optimizing on principal components directly, without using the DeepONet representation. B. Performance Summary For an incident wavelength of λ = 1100 nm and deflection angle θ = 60◦ , Figure 3 qualitatively presents six optimized metagrating patterns generated by NOTES. Panels (E) and (F) show rare cases with minor non-binarized regions, although most generated designs remain fully binary across all operating conditions. In nanophotonics, the aspect ratio, defined as the ratio between the device height and the minimum feature width, is an important indicator of fabrication difficulty, with larger values generally being harder to manufacture. The

6

A 8

0

96.41%

85.37%

30%

B 14

Training curve

max

mean

Seed 506 mean

12

96.41% max

99.19%

2 0

93.17% 94.56%

75%

Seed 658

mean 96.72%

Training curve

mean

10

0 20

max

99.19% max

30%

C

max

99.17%

2

mean 96.45%

0

75%

92.34% 93.59%

Seed 701 8

0 D

mean

max

99.17% max

30%

16

Training curve

98.18%

mean 96.35% 2

0

0 30%

98.18%

Fig. 5. Out-of-distribution performance of NOTES for visible-wavelength metagrating optimization. Panels (A–D) show optimized metagrating patterns at a wavelength of 530nm, which lies far outside the training range whose lower bound is 900nm. The mean (red dashed line) and maximum (green dashed line) values of the distribution are indicated in the histogram. The target deflection angles are (A) 40°, (B) 50°, (C) 60°, and (D) 70°.

obtained designs exhibit aspect ratios of 65.52, 16.38, 16.38, 16.38, 16.38, and 21.84. Overall, 95% of the designs generated by NOTES have aspect ratios below 33. Figure 4 presents the distribution of optimized designs for a wavelength of 1100 nm and a deflection angle of 60° for NOTES, CMA-ES, and L-BFGS + DeepONet. CMAES exhibits substantially higher variance in efficiency than NOTES. As shown in panels (A) and (C), most designs produced by NOTES converge to efficiencies higher than 92%, whereas CMA-ES yields a much broader distribution, with

75%

92.15% 94.22%

Fig. 6. This figure demonstrates the robustness of NOTES when the neural operator is trained with different random seeds. The mean (red dashed line) and maximum (green dashed line) values of the distribution are indicated in the histogram. The corresponding neural operator training curves are shown in insets.

most designs below 65% and only a few outliers near 94%. This behavior reflects the difficulty of applying evolutionary strategies directly in a high-dimensional design space. LBFGS + DeepONet also exhibits higher variance than NOTES, with a left-skewed efficiency distribution, suggesting frequent convergence to poor local optima, analyzed in more detail in the next section. These results indicate that CMA-ES is more effective than L-BFGS at identifying high-quality optima in the learned latent space. From the results in Table I, among all methods, NOTES demonstrates the best overall performance in most cases. Even when it does not achieve the top result, its performance remains within 2% of the best. NOTES consistently achieves the highest mean efficiencies and the lowest variance. Designs generated by NOTES also have a high probability of satisfying

7

the binarization constraint. 95% of the designs from NOTES have aspect ratios less than 33, and 97% of the results are fully binary. By learning from existing data, it generates new designs with efficiencies that are, on average, 3% higher than GLOnet [40]. Although this improvement may appear modest, it is quite significant given that GLOnet [40] has state-ofthe-art performance with many designs close to the global optimum, 100%. L-BFGS + DeepONet sometimes yields the highest-performing designs, but they are not necessarily binarized and, therefore, nonphysical. Baseline CMA-ES, which uses four times more iterations than NOTES, exhibits highly unstable performance, showing variance that ranges from the highest to nearly zero and efficiencies that vary from the best to the worst among all methods. For the MBB beam application, we find that NOTES achieves a compliance of 246, demonstrating a 1.6% improvement over L-BFGS-based topology optimization (SI section 2.4 & Table S1). NOTES also reaches the same best performance as L-BFGS + DeepONet. Regarding computational costs, solving the PDE is the most expensive part, largely dominating the DeepONet training costs. Each iteration, the evaluation of the objective function requires solving the PDE. We have found that NOTES has a faster convergence rate of up to one order of magnitude compared to direct CMA-ES (see Figure S1 in SI), achieving a better tradeoff between performance and the number of PDE evaluations.

the proposed nonlinear DeepONet decoder used in NOTES across both applications. For fairness, the PCA reconstruction is employed with the same post-processing transformations as NOTES in each application. For the 1D nanophotonic beam deflector, PCA-decoder + CMA-ES performs comparably to NOTES and in some cases achieves slightly higher best efficiencies, with an average improvement of approximately 1% under the same computational budget (Figure S6 in SI). However, in the more geometrically complex 2D MBB beam, the PCA-decoder substantially underperforms NOTES (Figure S14 in SI). Under the same budget, NOTES achieves a best compliance of 246, whereas PCA-decoder + CMA-ES only achieves approximately 394, and improves to approximately 300 with twice the computational budget. This contrast reveals that the effectiveness of linear PCA reconstruction strongly depends on the degree of dimensional compression and the complexity of the feasible topology manifold. In the 1D case, reducing from 256 dimensions to a 25-dimensional latent space still preserves enough geometric information for linear reconstruction, whereas in the 2D case the compression from 12,288 dimensions to only 60 latent variables requires recovering a highly nonlinear mapping that PCA cannot adequately represent. These results confirm that the performance gain of NOTES does not arise only from the PCA latent space, but also from the nonlinear expressive power of the DeepONet decoder, which reconstructs a more complex manifold of feasible highperforming PDE-constrained topologies.

C. L-BFGS + DeepONet vs NOTES Table I shows that L-BFGS + DeepONet exhibits substantially higher variance than NOTES in the photonic application. To understand the reason for such behavior, we compared the L2 gradient norms with respect to the latent space parameters (Figure S4 in SI). With a 25-dimensional latent space, the average gradient norm for NOTES is 2×10−4 , whereas the average gradient norm for L-BFGS + DeepONet is 7.5 × 10−3 , both achieving designs that are near local optima. To further validate this observation, we applied additional gradient descent steps to the solutions obtained from L-BFGS + DeepONet. Although the average gradient norm was further reduced to 4×10−3 , the efficiency distribution remained nearly unchanged (Figure S5 in SI). However, the significantly higher variance observed for L-BFGS + DeepONet suggests that the optimization process frequently converges to poor local optima. We infer from these results that the larger variance of L-BFGS + DeepONet is primarily caused by convergence to poor local optima. In contrast, for the MBB beam application, L-BFGS + DeepONet and NOTES achieve comparable performance and converge to similar designs (Figure S11 & S16 in SI). These results indicate that the latent optimization landscape presents fewer poor local optima. D. DeepONet vs PCA decoder To examine whether DeepONet is necessary to achieve strong performance, we conduct additional comparative studies between a linear PCA-decoder coupled with CMA-ES and

E. Reusability The ability of NOTES to generate high-performance designs beyond the training distribution is demonstrated in Table I. During data selection, only designs with pre-computed/labeled efficiency ≥ 0.9 were retained from the MetaNet dataset [42]. For the angle–wavelength combinations of (1100 nm, 40◦ ), (900 nm, 50◦ ), and (1100 nm, 60◦ ), no design with precomputed/labeled efficiency satisfying this threshold is available. However, after reevaluating designs in the training data, we find that for (1100 nm, 40◦ ), the best design achieves an efficiency of 75%, whereas NOTES produces a new design with 87% efficiency. For (900 nm, 50◦ ), the best design available in the training data achieves 98% efficiency; NOTES discovers designs reaching 99%, with smaller variance and a substantially faster convergence rate. For (1100 nm, 40◦ ), the best design available in the training data achieves 90% efficiency, while NOTES achieves 95%, representing a 5% gain. Beyond these three illustrations, NOTES consistently generates new designs that outperform the training data. Figure 5 further demonstrates the transferability of NOTES. By learning a compact neural reparameterization, NOTES accelerates optimization and generalizes effectively under unseen operating conditions. Panels (A–D) present metagrating optimizations at 530 nm (green), far outside the training wavelength range, targeting deflection angles of 40◦ , 50◦ , 60◦ , and 70◦ . Despite operating at an out-of-distribution wavelength, the method yields patterns with near-perfect deflection efficiency (≈ 100%). It is important to note that these are

8

simulation results, in which the efficiencies are overestimated because visible-light absorption in silicon is neglected. In panel (A), the run-to-run variance is high, while our usual empirical findings align more closely with panels (B–D). This highlights the occasional instability but strong reusability of the trained operator. From (A) to (D), the aspect ratios of those designs are low at 10.09, 15.03, 13.59, and 16.39 respectively, all of which remain within a reasonable fabrication range. For the MBB beam application, we also find that the trained DeepONet can be used for different resolutions and volume constraints. For example, we apply the trained DeepONet to optimize designs at the 96×32 resolution with a volume constraint of 0.5 and at the 384×128 resolution with a volume constraint of 0.3. In both cases, the pretrained NOTES perform well compared to the results of pixel-LBFGS from previous work [35]. At the 96×32 resolution, pretrained NOTES achieves a design that is within 2.4% of the best L-BFGS design (Figure S13 in SI). At the 384×128 resolution, pretrained NOTES is within 1.8% of the best L-BFGS design (Figure S12 in SI). Thanks to its meshless property, DeepONet seamlessly allows for different resolutions. These results demonstrate the transferability of the DeepONet representation.

F. Robustness of NOTES The stability analysis demonstrates the robustness of NOTES when the neural operator is trained with different random seeds (Figure 6). Despite variations in initializations, all neural operators produce nearly identical training curves. The resulting NOTES produce design distributions with similar shapes, means, and maximum efficiencies, indicating that the NOTES framework and induced latent space are robust to variations in neural-operator initialization.

G. Comparison Between Randomly-sampled and Physicsinformed Data To further validate that the performance of NOTES arises from its physics-informed training data rather than from the DeepONet architecture alone, we conduct an ablation study on both the nanophotonic and MBB beam applications using DeepONets trained on designs sampled uniformly at random instead of topology-optimized designs. For the nanophotonic application, although both methods achieve similar maximum efficiencies, the randomly sampled approach exhibits significantly larger variance in performance and produces more designs with high aspect ratios (Figure S7 in SI). For the MBB beam application, NOTES trained on randomly-sampled data generates designs with compliance 735 (Figure S15 in SI), far worse than the compliance of 246 achieved by its physics-informed counterpart. These results demonstrate that the key advantage of NOTES does not arise only from latentspace dimensionality reduction or the DeepONet decoder, but also from the physically meaningful manifold induced by high-performance PDE-optimized training data, which embeds transferable topological priors essential for efficient downstream search.

VII. D ISCUSSION Our framework imposes inductive biases in the neural reparameterization through the use of physics-informed designs, the latent-space representation, and the neural operator. This reparameterization scheme enables optimization to take place in a fully unconstrained space while satisfying binarization constraints. Currently, NOTES implements the binarization constraint through the neural operator. For more complex devices, rigorous methods have been developed to enforce fabrication constraints [38], [46], [47]. Most designs generated by the neural operator are already fully binary when coupled with CMA-ES. However, when coupled with gradient-based methods, the DeepONet often struggles to produce binarized designs. Although the exact cause requires further investigation, this difficulty may be dependent on the optimization formulation [48], [49]. The latent space induced by the neural operator does not inherently eliminate poor local optima. Consequently, gradient-based methods may be trapped in poor local minima, whereas CMA-ES is more robust and consistent in a low-dimensional space. Our physics-informed ablation study suggests that dimensionality reduction inherently restricts the feasible design manifold explored during optimization. Consequently, incorporating a physics-informed prior is critically important for guiding the search toward high-performance regions of the design space. The additional experiment with PCA decoder shows that the neural operator learns a better non-linear mapping to a more complex design manifold in high-dimensional space than a linear decoder. NOTES outperforms its counterparts in ablation studies for the 2D structural optimization problem, whereas the performance gap is smaller in the simpler 1D case, suggesting that the benefits from NOTES increase with the problem’s complexity. Although NOTES transfers effectively across inverse design conditions and parameters within the same PDE family, the limits of representation transferability across substantially different topology classes or physical regimes remain an open question. VIII. C ONCLUSION Overall, NOTES establishes an effective physics-informed evolutionary strategy for generating high-performance and physically realistic designs. By combining dimensionality reduction, neural operator representations, and CMA-ES, the framework achieves rapid and consistent convergence. Our results demonstrate that DeepONet is an effective nonlinear decoder that captures high-performance regions of topology manifolds. NOTES also demonstrates strong reusability across different resolutions, optimization constraints, and PDE parameters. Finally, we find that CMA-ES is more effective than L-BFGS at escaping poor local optima within the learned latent space. In this work, NOTES is demonstrated using PCA as the dimensionality reduction method. In future work, other latent-space representations may be explored, including neural encoder-based approaches [18], [50], [51]. Different latentspace representations may affect convergence stability and computational cost. Although gradient-based optimization

9

Angle / Wavelength (deg / nm) 40 / 900

Metric

Topology Opt.

Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std Mean Max Std

68% 94% 16% 65% 90% 14% 61% 87% 10% 64% 93% 16% 55% 95% 16% 49% 91% 10% 59% 93% 18% 56% 92% 14% 52% 79% 15% 59% 92% 13% 62% 84% 12% 59% 84% 14%

GLOnet

L-BFGS + DeepONet

CMA-ES

NOTES

88% 69% 86% 93% 95% 95% 94% 95% 8% 20% 16% 2% 40 / 1000 69% 72% 66% 88% 92% 91% 86% 91% 20% 16% 20% 2% 40 / 1100 60% 74% 82% 82% 86% 85% 88% 87% 2% 15% 9% 6% 50 / 900 90% 81% 92% 98% 98% 99% 96% 99% 1% 10% 17% 2% 50 / 1000 85% 81% 94% 96% 97% 96% 97% 96% 12% 17% 2% 1% 50 / 1100 77% 82% 91% 93% 91% 98% 94% 96% 11% 13% 1% 2% 60 / 900 73% 86% 99% 98% 97% 99% 99% 99% 13% 20% 0.2% 1% 60 / 1000 85% 84% 97% 98% 99% 98% 99% 98% 17% 16% 1% 0.5% 60 / 1100 59% 82% 60% 93% 80% 95% 94% 95% 17% 12% 18% 1.5% 70 / 900 83% 81% 89% 96% 98% 98% 99% 98% 2% 14% 17% 13% 70 / 1000 76% 78% 61% 97% 93% 99% 70% 99% 18% 18% 4% 1.5% 70 / 1100 65% 74% 61% 89% 90% 70% 90% 84% 14% 17% 5% 1.5% TABLE I P ERFORMANCE COMPARISON OF ALL METHODS . T HE BEST MEAN ( GREEN ), BEST MAXIMUM ( BLUE ), AND LOWEST STANDARD DEVIATION ( ORANGE ) ARE HIGHLIGHTED . A MONG ALL METHODS , NOTES ACHIEVES THE BEST PERFORMANCE , CONSISTENTLY PRODUCING HIGH - PERFORMANCE , FULLY BINARIZED DESIGNS WHILE EXHIBITING THE LOWEST VARIANCE , INDICATING RELIABLE CONVERGENCE TO HIGH - QUALITY OPTIMA .

coupled with neural representations (L-BFGS + DeepONet) achieves promising design performance and lower computational cost, the resulting designs are often not fully binary. Future work will explore subpixel-smoothed projection [52] as a hard binarization constraint to improve the physical realism of gradient-based designs while mitigating the illconditioning associated with the hard thresholding of fabrication constraints.

IX. DATA AND C ODE AVAILABILITY The results of this work used the Metanet [42] dataset. The implementation of the MBB beam application follows the same framework as the prior work [35], but instead of TensorFlow, we implement the code in PyTorch. The code implementation is available on GitHub.

X. ACKNOWLEDGMENT

X. Huang and R. Pestourie acknowledge support from the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, U.S. National Science Foundation under Grants No. IIS 2435905, and National Institute of Biomedical Imaging and Bioengineering of the National Institutes of Health under Award Number R21EB036343. G. Zhang acknowledges support from the U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research, Applied Mathematics program, under the contracts ERKJ388 and ERKJ456 at Oak Ridge National Laboratory. L.Lu was supported by the U.S. DOE ASCR under Grants No. DE-SC0025593 and No. DESC0025592, and the U.S. National Science Foundation under Grants No. DMS-2347833 and No. DMS-2527294.

10

R EFERENCES [1] Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via deeponet based on the universal approximation theorem of operators. Nature machine intelligence, 3(3):218–229, 2021. [2] Nikolaus Hansen. The cma evolution strategy: A tutorial. arXiv preprint arXiv:1604.00772, 2016. [3] M. Raissi, P. Perdikaris, and G.E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019. [4] Lu Lu, Raphael Pestourie, Wenjie Yao, Zhicheng Wang, Francesc Verdugo, and Steven G Johnson. Physics-informed neural networks with hard constraints for inverse design. SIAM Journal on Scientific Computing, 43(6):B1105–B1132, 2021. [5] Nikolaus Hansen and Andreas Ostermeier. Completely derandomized self-adaptation in evolution strategies. Evolutionary computation, 9(2):159–195, 2001. [6] Anne Auger and Nikolaus Hansen. A restart cma evolution strategy with increasing population size. In 2005 IEEE congress on evolutionary computation, volume 2, pages 1769–1776. IEEE, 2005. [7] Nikolaus Hansen. The CMA Evolution Strategy: A Comparing Review, pages 75–102. Springer Berlin Heidelberg, Berlin, Heidelberg, 2006. [8] Youhei Akimoto and Nikolaus Hansen. Projection-based restricted covariance matrix adaptation for high dimension. In Proceedings of the Genetic and Evolutionary Computation Conference 2016, pages 197– 204, 2016. [9] Ilya Loshchilov, Tobias Glasmachers, and Hans-Georg Beyer. Large scale black-box optimization by limited-memory matrix adaptation. IEEE Transactions on Evolutionary Computation, 23(2):353–358, 2018. [10] Adam Slowik and Halina Kwasnicka. Evolutionary algorithms and their applications to engineering problems. Neural Computing and Applications, 32(16):12363–12379, 2020. [11] Richard E Bellman and Stuart E Dreyfus. Applied dynamic programming. Princeton university press, 2015. [12] Kento Uchida, Teppei Yamaguchi, and Shinichi Shirakawa. Covariance matrix adaptation evolution strategy for low effective dimensionality. arXiv preprint arXiv:2412.01156, 2024. [13] Gawel Kus and Miguel A Bessa. Gradient-free neural topology optimization: towards effective fracture-resistant designs. Computational Mechanics, 75(4):1327–1355, 2025. [14] Hongshu Guo, Wenjie Qiu, Zeyuan Ma, Xinglin Zhang, Jun Zhang, and Yue-Jiao Gong. Advancing cma-es with learning-based cooperative coevolution for scalable optimization. arXiv preprint arXiv:2504.17578, 2025. [15] Michael Hinze, René Pinnau, Michael Ulbrich, and Stefan Ulbrich. Optimization with PDE constraints, volume 23. Springer Science & Business Media, 2008. [16] Rebekka V Woldseth, Niels Aage, J Andreas Bærentzen, and Ole Sigmund. On the use of artificial neural networks in topology optimisation. arXiv preprint arXiv:2208.02563, 2022. [17] Wei Ma, Feng Cheng, Yihao Xu, Qinlong Wen, and Yongmin Liu. Probabilistic representation and inverse design of metamaterials based on a deep generative model with semi-supervised learning strategy. Advanced Materials, 31(35):1901111, 2019. [18] Reza Marzban, Hamed Abiri, Raphaël Pestourie, and Ali Adibi. Hilab: A hybrid inverse-design framework. Small Methods, page e00975, 2025. [19] Zhaocheng Liu, Dayu Zhu, Sean P Rodrigues, Kyu-Tae Lee, and Wenshan Cai. Generative model for the inverse design of metasurfaces. Nano letters, 18(10):6570–6576, 2018. [20] Sudhanshu Singh, Rahul Kumar, Soumyashree S Panda, and Ravi S Hegde. Deep-learning enabled photonic nanostructure discovery in arbitrarily large shape sets via linked latent space representation learning. Digital Discovery, 3(8):1612–1623, 2024. [21] Sudhanshu Singh, Rahul Kumar, Praveen Singh, and Ravi Hegde. Active learning for efficient nanophotonics inverse design in large and diverse design spaces. Optics Express, 33(10):20308–20321, 2025. [22] Jian Zhang, Jin Yuan, Chuanzhen Li, and Bin Li. An inverse design framework for isotropic metasurfaces based on representation learning. Electronics, 11(12):1844, 2022. [23] Keon Ko, Min Woo Cho, Kyungjun Song, Dong Yong Park, and Sang Min Park. Inverse design of non-parametric acoustic metamaterials via transfer-learned dual variational autoencoder with latent space-based data augmentation. Engineering Applications of Artificial Intelligence, 151:110735, 2025.

[24] Zekun Ren, Siyu Isaac Parker Tian, Juhwan Noh, Felipe Oviedo, Guangzong Xing, Jiali Li, Qiaohao Liang, Ruiming Zhu, Armin G Aberle, Shijing Sun, et al. An invertible crystallographic representation for general inverse design of inorganic crystals with targeted properties. Matter, 5(1):314–335, 2022. [25] Ning Miao, Emile Mathieu, N Siddharth, Yee Whye Teh, and Tom Rainforth. On incorporating inductive biases into vaes. arXiv preprint arXiv:2106.13746, 2021. [26] Yaniv Yacoby, Weiwei Pan, and Finale Doshi-Velez. Failure modes of variational autoencoders and their effects on downstream tasks. arXiv preprint arXiv:2007.07124, 2020. [27] Tinghao Guo, Danny J Lohan, Ruijin Cang, Max Yi Ren, and James T Allison. An indirect design representation for topology optimization using variational autoencoder and style transfer. In 2018 AIAA/ASCE/AHS/ASC structures, structural dynamics, and materials conference, page 0804, 2018. [28] Pengzhan Jin, Shuai Meng, and Lu Lu. Mionet: Learning multiple-input operators via tensor product. SIAM Journal on Scientific Computing, 44(6):A3490–A3514, 2022. [29] Lu Lu, Xuhui Meng, Shengze Cai, Zhiping Mao, Somdatta Goswami, Zhongqiang Zhang, and George Em Karniadakis. A comprehensive and fair comparison of two neural operators (with practical extensions) based on fair data. Computer Methods in Applied Mechanics and Engineering, 393:114778, 2022. [30] Anran Jiao, Haiyang He, Rishikesh Ranade, Jay Pathak, and Lu Lu. One-shot learning for solution operators of partial differential equations. Nature Communications, 16(1):8386, 2025. [31] Anran Jiao, Qile Yan, Jhn Harlim, and Lu Lu. Solving forward and inverse pde problems on unknown manifolds via physics-informed neural operators. arXiv preprint arXiv:2407.05477, 2024. [32] Min Zhu, Handi Zhang, Anran Jiao, George Em Karniadakis, and Lu Lu. Reliable extrapolation of deep neural operators informed by physics or sparse observations. Computer Methods in Applied Mechanics and Engineering, 412:116064, 2023. [33] Minglang Yin, Nicolas Charon, Ryan Brody, Lu Lu, Natalia Trayanova, and Mauro Maggioni. A scalable framework for learning the geometrydependent solution operators of partial differential equations. Nature Computational Science, 4(12):928–940, 2024. [34] Chenhui Kou, Yuhui Yin, Min Zhu, Shengkun Jia, Yiqing Luo, Xigang Yuana, and Lu Lu. Efficient neural topology optimization via active learning for enhancing turbulent mass transfer in fluid channels. arXiv preprint arXiv:2503.03997, 2025. [35] Stephan Hoyer, Jascha Sohl-Dickstein, and Sam Greydanus. Neural reparameterization improves structural optimization. arXiv preprint arXiv:1909.04240, 2019. [36] Archis S Joglekar. Generative neural reparameterization for differentiable pde-constrained optimization. arXiv preprint arXiv:2410.12683, 2024. [37] Lu Lu, Raphaël Pestourie, Steven G Johnson, and Giuseppe Romano. Multifidelity deep neural operators for efficient learning of partial differential equations with application to fast inverse design of nanoscale heat transport. Physical Review Research, 4(2):023210, 2022. [38] Mo Chen, Rasmus E Christiansen, Jonathan A Fan, Göktuğ Işiklar, Jiaqi Jiang, Steven G Johnson, Wenchao Ma, Owen D Miller, Ardavan Oskooi, Martin F Schubert, et al. Validation and characterization of algorithms and software for photonics inverse design. Journal of the Optical Society of America B, 41(2):A161–A176, 2024. [39] Jing Liu, Ruhul Sarker, Saber Elsayed, Daryl Essam, and Nurhadi Siswanto. Large-scale evolutionary optimization: A review and comparative study. Swarm and Evolutionary Computation, 85:101466, 2024. [40] Jiaqi Jiang and Jonathan A Fan. Simulator-based training of generative neural networks for the inverse design of metasurfaces. Nanophotonics, 9(5):1059–1069, 2020. [41] Yongha Kim, Anthony W Jung, Sanmun Kim, Kevin Octavian, Doyoung Heo, Chaejin Park, Jeongmin Shin, Sunghyun Nam, Chanhyung Park, Juho Park, et al. Meent: Differentiable electromagnetic simulator for machine learning. arXiv preprint arXiv:2406.12904, 2024. [42] Jiaqi Jiang, Robert Lupoiu, Evan W Wang, David Sell, Jean Paul Hugonin, Philippe Lalanne, and Jonathan A Fan. Metanet: a new paradigm for data sharing in photonics research. Optics express, 28(9):13670–13681, 2020. Dataset available online at MetaNet. Accessed: Nov. 11, 2025. [43] Christopher M Bishop and Nasser M Nasrabadi. Pattern recognition and machine learning, volume 4. Springer, 2006. [44] Lu Lu, Xuhui Meng, Zhiping Mao, and George Em Karniadakis. DeepXDE: A deep learning library for solving differential equations. SIAM Review, 63(1):208–228, 2021.

11

[45] Nikolaus Hansen, Y Akimoto, and P Baudis. Cma-es/pycma: r4. 0.0, 2024. [46] Pavel Terekhov, Shengyuan Chang, Md Tarek Rahman, Sadman Shafi, Hyun-Ju Ahn, Linghan Zhao, and Xingjie Ni. Enhancing metasurface fabricability through minimum feature size enforcement. Nanophotonics, 13(17):3147–3154, 2024. [47] Jacob M Hiesener, C Alex Kaylor, Joshua J Wong, Prankush Agarwal, and Stephen E Ralph. Seeded topology optimization for commercial foundry integrated photonics. arXiv preprint arXiv:2503.00199, 2025. [48] Dimitri P Bertsekas. Nonlinear programming. Journal of the Operational Research Society, 48(3):334–334, 1997. [49] Jorge Nocedal and Stephen J Wright. Numerical optimization. Springer, 2006. [50] Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8):1798–1828, 2013. [51] Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436–444, 2015. [52] Alec M Hammond, Ardavan Oskooi, Ian M Hammond, Mo Chen, Stephen E Ralph, and Steven G Johnson. Unifying and accelerating level-set and density-based topology optimization by subpixel-smoothed projection. Optics Express, 33(16):33620–33642, 2025.

Supplementary Information

In the following sections, we provide additional figures, ablation studies, and reproducible experimental details for NOTES applied to both the nanophotonic beam deflector problem and the MBB beam structural optimization problem. The nanophotonic application section (Section 1) includes convergence-rate comparisons with standard CMA-ES, DeepONet training curves, reconstruction quality, latent-space gradient analyses, additional optimization experiments, and ablation studies comparing DeepONet and PCA-decoder as well as physics-informed and randomly sampled training data. The MBB beam application section (Section 2) presents reproducible details for data generation and processing, PCA dimensionality reduction, DeepONet architecture and training, optimization settings, transfer-learning experiments across resolutions and volume fractions, and additional ablation studies comparing NOTES against PCA-decoder, randomly sampled training data, and L-BFGS + DeepONet baseline. These supplementary results further support the robustness and reproducibility of the proposed NOTES framework.

1

Additional Details for Nanophotonic application

1.1

Convergence rate comparison

Figure S1 shows the comparison between NOTES and standard CMA-ES. NOTES optimizes over a 25-dimensional space, while standard CMA-ES optimizes over a 256-dimensional space. Each method applies 30 runs, and the bestso-far objective values are plotted against the number of function evaluations. We observe that NOTES achieves a faster convergence rate than CMA-ES due to the reduced degrees of freedom.

(a) Example 1

(b) Example 2

Figure S1: Comparison of convergence histories of NOTES (orange) and standard CMA-ES (blue). The in-figure legend entry “Baseline” refers to standard CMA-ES. The threshold, shown as a dashed gray horizontal line, is defined as one standard deviation above the mean value of standard CMA-ES at the last epoch. The curves represent mean value at each epoch. The upper and lower bounds indicate one standard deviation above and below the mean, respectively. Each dot on the curves corresponds to an interval of 10 epochs.

1.2

Training curve

Figure S2 shows the loss curve of DeepONet trained for 400,000 epochs. Every 40,000 epochs, the scale factor of the sigmoid function in the DeepONet is doubled. Although the gap between the training loss and validation loss

1

increases over the epochs, the validation metric, mean squared error (MSE), remains approximately unchanged at around 0.06.

Figure S2: Training loss is shown by the blue curve, while validation loss is shown by the orange curve. An additional validation metric, mean squared error (MSE), is shown by the green curve.

2

1.3

Reconstruction Quality

(a) Sample 1

(b) Sample 2

(c) Sample 3

(d) Sample 4

Figure S3: Visualization of the reconstruction quality of DeepONet for the nanophotonic application. The upper panel shows the ground-truth design (black), while the lower panel shows the corresponding design reconstructed by DeepONet (red). Figure S3 shows the reconstruction quality of DeepONet. Overall, DeepONet predicts the locations and widths of the ridges in each design quite well, but some ridge edges are not sharp.

1.4

Additional Gradient Descent

To verify whether the solutions produced by L-BFGS + DeepONet are trapped near local optima, we performed an additional gradient descent refinement experiment. Specifically, we apply 2,000 iterations of gradient descent with a learning rate of 10−3 to the latent variables initialized from the L-BFGS + DeepONet solutions. The learning rate was selected based on preliminary gradient-norm-versus-iteration tests, in which larger learning rates resulted in unstable optimization trajectories. If the L-BFGS + DeepONet solutions are not near stationary points, this additional refinement should substantially reduce both the latent gradient norms and the corresponding objective values. In this experiment, we considered 100 optimized designs at a deflection angle of 40◦ and a wavelength of 1000 nm, and compared the latent-space gradient norms before and after refinement.

3

(a) CDF of Gradient Norm

(b) Bar plot of Gradient Norm

Figure S4: Comparison of latent-space gradient norms for NOTES, L-BFGS + DeepONet, and L-BFGS + DeepONet followed by gradient descent. The comparison of latent-space gradient norms (Figure S4) shows that additional gradient descent consistently reduces the gradient magnitude of the L-BFGS + DeepONet solutions. The average gradient norm decreases from 7.5 × 10−3 before gradient descent to 4.0 × 10−3 after gradient descent, while the standard deviation decreases from 1.41 × 10−2 to 6.3 × 10−3 .

(a) CDF of Efficiency

(b) Bar plot of Efficiency

Figure S5: Comparison of the final objective values for designs generated by NOTES, L-BFGS + DeepONet, and L-BFGS + DeepONet followed by gradient descent. The curves for L-BFGS + DeepONet and L-BFGS + DeepONet followed by gradient descent overlap almost completely. In contrast, the comparison of deflection efficiency (Figure S5) shows that the efficiency distribution remains nearly unchanged before and after the additional gradient descent refinement. Although the latent-space gradient norms are reduced, the objective values exhibit no meaningful improvement. This indicates that the additional gradient descent steps only move the solutions within the same attraction basin without escaping to significantly better designs. Therefore, the solutions produced by L-BFGS + DeepONet are already trapped near poor local optima in the latent landscape. 4

1.5

DeepONet vs PCA-decoder

For this ablation study, we replace the DeepONet decoder in NOTES with a linear PCA decoder for the 1D nanophotonic application. We use the same dataset employed to train the DeepONet, consisting of all designs with precomputed deflection efficiencies greater than 0.9, to derive the PCA representation. Specifically, PCA decomposition is applied to obtain the first 25 principal components together with the mean vector. The PCA decoder reconstructs the designs by reversing the PCA projection using these retained components.

σ(x; β, η) =

1 1 + exp(−βx − η)

(1)

To enforce optimization constraints, we apply the same custom sigmoid function 1 used in NOTES, with a scale factor of 10 and a threshold of 0. The optimization framework is otherwise identical to NOTES, except that the DeepONet decoder is replaced by the PCA decoder. CMA-ES is initialized from a unit Gaussian distribution in the latent space with an initial step size of σ0 = 2.0. As shown in Figure S6, CMA-ES combined with the PCA decoder achieves slightly better performance than NOTES. Both DeepONet and PCA-decoder are consistent at generating high-performance designs.

Figure S6: Comparison of the final efficiency for the PCA decoder and DeepONet, both combined with CMA-ES, at a wavelength of 1100 and an angle of 60.

1.6

Randomly-Sampled vs Physics-Informed data

To further validate that the performance of NOTES arises from its physics-informed training data rather than from the DeepONet architecture alone, we conduct an ablation study on the nanophotonic application using randomly sampled designs. Specifically, we generate a dataset of the same size as the physics-informed training dataset by independently sampling each pixel from a Bernoulli distribution. A new DeepONet is trained on this randomly sampled dataset using exactly the same training procedure and architecture as the physics-informed counterpart. The results indicate that NOTES trained on randomly generated designs performs worse than NOTES trained on physics-informed designs. Although both methods achieve similar maximum efficiencies (approximately 94–95%), the randomly initialized approach exhibits significantly larger variance in performance. The variance in efficiency is approximately one order of magnitude larger than that of the physics-informed counterpart. In addition, more than 75% of the designs generated from random initialization have aspect ratios greater than 33, whereas fewer than 5% of the physics-informed designs exceed this threshold. These results indicate that the physics-informed prior substantially improves optimization consistency. Consequently, obtaining high-performance designs using randomly sampled data would require significantly higher computational cost.

5

(a) Distribution of final efficiency for physics-informed data

(b) CDF of aspect ratios for physics-informed data

(c) Distribution of final efficiency for randomly-sampled data

(d) CDF of aspect ratios for randomly-sampled data

Figure S7: Comparison between physics-informed and randomly sampled training data. The red dashed vertical line in the CDF plots highlights the threshold used to distinguish designs with low and high aspect ratios.

2

Experiment details for Structural Optimization

The second application is a structural optimization problem. We apply NOTES to Messerschmitt-Bölkow-Blohm (MBB) beams at a resolution of 192 × 64, which corresponds to 12,288 pixels. Using PCA, the latent space is reduced to 60 dimensions. By training DeepONet with augmented data, NOTES can discover high-performance designs. Moreover, when the trained DeepONet is transferred to optimize MBB beams with different resolutions and volume constraints, the method still produces competitive results.

2.1

Data Generation

The training data is generated using topology optimization with L-BFGS optimizer. The volume constraint is enforced using a custom sigmoid function with a threshold derived via binary search. The initial design is set such that all entries are equal to the volume constraint, a predefined value between 0 and 1. After optimization is completed, we randomly erode the optimized design for 1 to 8 iterations. We repeat this process to generate 1,000 designs.

6

2.2

Data Processing

The median compliance of the designs is about 270. We generate designs with compliance greater than 250 as training data. There are a total of 857 such designs. For each design, we apply dilation and erosion to obtain two augmented designs. In total, there are 2056 designs for training and 515 designs for validation. Figure S8 shows designs after dilation and erosion. We apply PCA and keep the first 60 principal components to form the latent space. According to Figure S9a, these components explain about 83% of the variance in the designs. Qualitatively, the dataset projected onto the first and second principal components shows clear clustering (Figure S9b), indicating that the dataset may have exploitable low dimensionality. Before the training of DeepONet, the PCA coordinates are normalized to have zero-mean and unit standard deviation along each dimension. This additional step simplifies the selection of the sampling radius for the downstream CMA-ES optimization.

Figure S8: Examples of eroded and dilated designs.

7

(a) PCA explained variance

(b) Visualization of designs in the first two principal components

Figure S9: PCA analysis of the eroded, dilated, and original designs.

2.3

Model Training

A DeepONet with 1.6M parameters is trained using this training data. Both the trunk and branch networks consist of four fully connected hidden layers with 512 neurons per layer and ReLU activation functions. The network weights are initialized using Glorot normal initialization. We apply a custom sigmoid function after the last layer of DeepONet to ensure that the outputs are between 0 and 1. The loss function is binary cross-entropy loss, with an additional binary soft-constraint term added to the objective. DeepONet is trained for 30K epochs. Figure S10 shows several comparisons between the true designs and the predicted designs. The reconstruction quality demonstrates that DeepONet captures large geometric features of the designs quite well.

8

Figure S10: Reconstruction quality of DeepONet for the 2D structural optimization application. The left column shows the eroded, dilated, and original designs, while the right column shows the corresponding designs reconstructed by DeepONet.

2.4

Optimization

During optimization, we replace the custom sigmoid with the hard constraint transformation mentioned in main text (Section V-C). We apply NOTES with a computational budget capped at 2.4k objective function evaluations. The population size is set to 30, and the initial sampling radius is set to 2.0. The best design has a compliance of 246. It is worth noting that the training data consisted of segmented designs with compliance greater than 250. Yet, DeepONet learns geometric features that achieve a design that is 1.6% better than the best design in the training data.

Figure S11: Optimization results generated by NOTES. As Figure S11 shows, we observe mode collapse as most optimized designs look very similar. It may be due to an imbalance of geometric features of the training data. Many designs in the training data share the same shape at the center and have an L-shape support on the wing. We reuse the pretrained DeepONet to optimize designs at a coarser resolution of 96 × 32 with a volume constraint of 0.5, or at a finer resolution of 384×128 with a volume constraint of 0.3. As shown in Figure S13 and Figure S12, we observe competitive performance without retraining. Table S1 summarizes the optimization results. It indicates that the one-shot multifidelity approach is particularly beneficial for the finer resolution, where creating a dataset would have been one or two orders of magnitude more computationally expensive, while having slightly worse performance than results of pixel L-BFGS baseline from prior work [1]. 9

Best design Improvement against pixel L-BFGS

96 × 32, 0.5 214 -2.4%

192 × 64, 0.4 246 1.6%

384 × 128, 0.3 326 -1.8%

Table S1: Optimization Results

Figure S12: Transfer learning for higher resolution, from 192 × 64 to 384 × 128.

Figure S13: Transfer learning for lower resolution, from 192 × 64 to 96 × 32.

2.5

DeepONet vs PCA-decoder for 2D

For the PCA-decoder ablation study on the MBB beam application, we use the same training dataset for DeepONet. The PCA projection matrix and mean vector are obtained through flattened designs. The PCA decoder reconstructs the designs by reversing the PCA projection, adding the mean vector, and reshaping back to original shape. It replaces the DeepONet decoder within the NOTES optimization framework. CMA-ES is initialized from a unit Gaussian distribution in the latent space with an initial step size σ0 = 2.0. For the geometrically more complex 2D structural optimization problem, the PCA decoder substantially underperforms the proposed NOTES framework, as shown in Figure S14. Under the same computational budget, NOTES identifies designs with compliance as low as 246 for the 192 × 64 MBB beam case, whereas CMA-ES coupled with the PCA decoder only achieves a best compliance of approximately 394. Even when the PCA-decoder baseline is allocated roughly twice the computational budget of NOTES, the best compliance only improves to around 300, after which no meaningful further improvement is observed. These results suggest that the nonlinear design manifold learned by DeepONet is significantly more expressive than a linear PCA representation for high-dimensional PDE-constrained structural optimization.

(a) Design produced by CMA-ES + PCA with the same com- (b) Design produced by CMA-ES + PCA with doubled computational cost putational cost Figure S14: Comparison of designs generated by CMA-ES + PCA under different computational budgets. 10

2.6

Randomly-sampled vs Physics-informed data in 2D

We further investigate the importance of the physics-informed component in NOTES through an ablation study on the MBB beam structural optimization problem. We generate random design logits with the same spatial resolution and sample size as the L-BFGS-generated dataset. Each sample is then processed using the volume constraint function, followed by cone filtering to enforce the prescribed volume fraction and to eliminate isolated small-scale features. This yields a total of 2,571 training samples for the ablation study. A DeepONet with the identical architecture used in NOTES is then trained on this synthetic dataset. At convergence, the binary cross-entropy loss remains at 0.675. For comparison, DeepONet trained with physicsinformed designs has a binary cross-entropy loss around 0.1. We then apply this trained DeepONet with CMA-ES to optimize with the same procedure as NOTES. In the best run, the initial compliance is 2466. After 80 iterations and 2400 function evaluations, the resulting compliance is 735, as shown in Figure S15.

Figure S15: Example design when NOTES is not physics-informed

2.7

L-BFGS + DeepONet vs NOTES in 2D

We apply L-BFGS + DeepONet to the 2D MBB beam structural optimization problem. Unlike the nanophotonic application, the volume constraint is explicitly enforced within the optimization pipeline. Consequently, all generated designs with L-BFGS optimizer satisfy the prescribed volume constraint. The same DeepONet used in NOTES is coupled with L-BFGS optimizer in this experiment. For each run, the initial starting point in the latent space is uniformly sampled along each dimension between -1 and 1. The initial learning rate is set to 0.5, which was selected after testing learning rates ranging from 0.001 to 1.

(a) Example 1

(b) Example 2

Figure S16: Optimization results generated by L-BFGS + DeepONet. As shown in Figure S16, NOTES and L-BFGS + DeepONet achieve comparable performance. Both methods converge to designs with similar structures and identical compliance values. However, L-BFGS + DeepONet requires approximately 10 times fewer function evaluations, making it a highly competitive alternative for this optimization problem.

References [1] Stephan Hoyer, Jascha Sohl-Dickstein, and Sam Greydanus. Neural reparameterization improves structural optimization. arXiv preprint arXiv:1909.04240, 2019.

11

Record · ID 349614 · SHA-256 9dcce19e5a6f805e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.