Bridging Spherical Black-Box Optimizers
Johannes Ackermann 1 Stefano Peluchetti 2
arXiv:2606.25761v1 [cs.LG] 24 Jun 2026
Abstract When gradient information is unavailable, blackbox optimization (BBO) methods provide a practical alternative. While Evolution Strategies (ES), Consensus-Based Optimization (CBO), Optimization via Integration (OVI), and related methods have each been studied independently, their connections remain underexplored. We unify these approaches within a common theoretical framework, revealing that they differ primarily in two design choices: fitness aggregation (controlling sharpness preference) and consensus scope (controlling modality). Leveraging these insights, we introduce hybrid optimizers that interpolate between existing methods. Our ES-OVI hybrid allows explicit control over the preference for flat minima, enabling a trade-off between performance and robustness in continuous control tasks. Our CBO-OVI hybrids combine the higherdimensional efficiency of parametric methods with the multimodal capabilities of particle-based approaches, achieving competitive results on language model merging under limited evaluation budgets. We validate our methods on standard BBO benchmarks and higher-dimensional locomotion tasks, demonstrating that the hybrid methods can outperform their constituent algorithms.
Parametric
Nonparametric
unimodal, work well in d ≫ 1
work well in d ≲ 30
ES
CBO unimodal
OVI λ = 1, d ≫ 1 mic,t
∗ θES
∗ θOVI
ES-OVI (Section 4.4) control convergence
SchedPol AdaPol (Section 4.5) multimodal in d ≫ 1
cCBO/pCBO multimodal
Figure 1. We investigate connections between the parametric ES, OVI, the nonparametric CBO, and further related optimizers. By utilizing these connections, we can derive hybrid methods, indicated by green arrows, that combine the strengths of existing optimizers: ES-OVI allows us to control convergence characteristics, SchedPol and AdaPol combine CBO and OVI updates, allowing us to obtain multiple optima in higher-dimensional tasks.
There are two main families of black-box optimizers, which are usually treated separately: Parametric and nonparametric methods. On the parametric side, a prominent family are Evolution Strategies (ES) (Rechenberg, 1973). Within ES, we specifically focus on Natural ES (NES) (Wierstra et al., 2008), which maintain a parametric search distribution, updating parameters based on the expectation of the objective over sampled points. Optimization via Integration (OVI) (Andrieu et al., 2024) is also parametric, but proposes to instead repeatedly tilt a search distribution by the objective values. On the non-parametric side, particlebased approaches such as Consensus Based Optimization (CBO) (Pinnau et al., 2017) evolve a population of candidate solutions. This is done based on the function value of all particles, or in the case of Polarized CBO (pCBO) (Bungert et al., 2025) only based on the value of nearby particles, allowing us to obtain multiple optima.
1. Introduction Stochastic gradient descent and its variants have become the dominant optimizers for deep learning models. However, gradients may be unavailable or not useful for training (Metz et al., 2022), for example when optimizing through simulators, using external APIs, or in Reinforcement Learning. In such cases, zeroth-order or black-box optimizers offer a solution as they only require objective function evaluations.
While ES, OVI, and CBO have each been linked to gradient descent (Salimans et al., 2017; Riedl et al., 2023; Andrieu et al., 2024), the connections between distribution-update methods (ES/OVI) and CBO particle methods have, to our knowledge, not been investigated. In this work, we propose a unified framework (Section 2) for these optimizers: a single master update equation (MU) from which ES, OVI, CBO, and their variants emerge as special cases, differing primarily in their fitness aggregation and interaction scope
JA did his work while interning at Sakana AI. 1 The University of Tokyo 2 Sakana AI. Correspondence to: Johannes Ackermann <[email protected]>. Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
1
Bridging Spherical Black-Box Optimizers Table 1. Instantiations of the master update (MU), unifying all considered BBO methods.
Method
Ψ(F ) (sharpness)
Kti,j (scope / interaction)
(µt , λt ) (transport)
σt s(x − m) (noise)
CH / OVI
1 (global)
(0, 1)
σ
ES
exp{−βF } η 1 − σ2 F − F
1 (global)
(0, 1)
σ
CBO pCBO
exp{−βF } exp{−βF }
(1 − λ, λ) (1 − λ, λ)
σ∥x − m∥ σ∥x − mi ∥
cCBO‡
exp{−βF }
(1 − λ, λ)
σ∥x − mi ∥
DE
exp{−βF }
1 (global) exp −∥xit − xjt ∥2 /(2κ2 ) j,c PNC pi,c t pt c=1 Ztc √ exp −∥xit − αt xjt ∥2 /(2(1 − αt ))
(µDDIM , λDDIM ) t t
σtDDIM
†
ES† : with antithetic sampling; Ψ depends on F as well (population mean of loss values). PN j,c j c cCBO‡ : pi,c t are cCBO assignments (4), Zt = j=1 pt exp{−βF (xt )}.
not depend on i, and we obtain a global consensus point mt . Most methods satisfy µt + λt = 1 (convex transport), but we keep (µt , λt ) separate to accommodate the DDIM (Song et al., 2021a) sampler in DE. If µt = 0 and λt = 1, then (MU) resamples the population around the consensus (xit+1 = mit + σt ϵit ), covering OVI/CH and ES.
(Table 1). We then leverage these insights to derive new optimizers that address the weaknesses of each method. We provide an overview in Figure 1. We empirically evaluate their convergence characteristics, and evaluate their performance on low-dimensional BBO benchmarks and higherdimensional locomotion tasks. Finally, we investigate their application to evolutionary model merging under limited evaluation budgets. We propose to address this setting as a multi-modal optimization problem, in which our proposed CBO-OVI combinations are successful.
Table 1 shows how the discussed BBO methods instantiate (MU). For ES, we assume mean-zero perturbations via antithetic sampling, a standard practice (Salimans et al., 2017). Under this assumption, ES can be written exactly in terms of (MU). Without it, the correspondence remains a close approximation for N ≫ 1. We defer to Appendix A.6 for details. The equivalence of CH and OVI is established in Appendix A.5. Finally, cCBO fits (MU) by using a clusterinduced interaction matrix Kt (Appendix A.7).
2. Unifying Framework We can view many black-box optimizers as iterating two steps: (i) compute a (possibly local) fitness-weighted consensus point, and (ii) update particles by (partially) moving toward this consensus while injecting exploration noise. This yields a single master update that recovers ES, OVI, CH, CBO, pCBO, clustered CBO (cCBO), and Diffusion Evolution (DE).
3. Background We consider a loss function F : Rd → R, which we aim to minimize. In the black-box optimization setting, we can evaluate F at any point, but cannot access its gradient ∇F .
d Master Update (MU) Given particles x1t , . . . , xN t ∈R , define
xit+1 = µt xit + λt mit + σt s(xit − mit )ϵit , mit =
N X
wti,j xjt ,
j=1
3.1. Parametric Methods
i i.i.d.
ϵt ∼ N (0, I),
Stochastic Smoothing (SS) constructs a differentiable surrogate for non-differentiable or black-box functions (Katkovnik & Kulchitsky, 1972; Spall, 2003; Nesterov & Spokoiny, 2017). Given a smoothing distribution with density π, the smoothed objective is JπSS (θ) = Eϵ∼π [F (θ + ϵ)]. Using the likelihood-ratio estimator, the gradient can be expressed as ∇θ JπSS (θ) = Eϵ∼π [F (θ + ϵ)∇ϵ − log π(ϵ)] (see Appendix A.2 for details). For a spherical Gaussian π = N (0, σ 2 I), this yields the gradient descent updates
ai,j wti,j = PN t i,ℓ , ℓ=1 at
j i,j ai,j t = Ψ(F (xt ))Kt .
(MU) Here Ψ is a fitness transformation (sharpness preference), Kti,j ≥ 0 is an interaction matrix controlling how particles influence each other (global vs. local vs. clustered interactions), λt controls attraction toward consensus, µt controls persistence, σt controls exploration, and s(·) determines noise scaling.
N
θt+1 = θt −
In all methods aside from cCBO, Kti,j is constructed from a kernel: Kti,j = kt (xit , xjt ). If Kti,j = 1, then mit does
η X F (θt + σϵi )ϵit , σN i=1
i.i.d.
ϵit ∼ N (0, I), (SS)
where η > 0 is the learning rate. 2
Bridging Spherical Black-Box Optimizers
Natural Evolution Strategies (NES) (Wierstra et al., 2008; 2014) optimize a parametric search distribution π(x; θ) to maximize expected fitness JπES (θ) = Ex∼π(·;θ) [−F (x)]. For a Gaussian π = N (θ, σ 2 I) with fixed variance, the NES updates are θt+1 = θt − ηEϵ∼N (0,σ2 I) [F (θt + ϵ)ϵ] .
particle toward the consensus, while the diffusion term adds exploration noise scaled by the Euclidean distance to the consensus. The Euler-Maruyama discretization of (1) with step size ∆t = 1 yields the CBO updates: xit+1 ← xit −λ(xit −mt )+σ∥xit −mt ∥ϵit ,
For spherical Gaussians, NES and SS yield equivalent updates (Section 4.1); we refer to both as Evolution Strategies (ES). Salimans et al. (2017) demonstrate the effectiveness of ES for reinforcement learning, showing favorable scaling with parallel computation. In practice, fitness shaping transformations (e.g., rank-based) are applied before computing updates, which we discuss more in Appendix A.1.
Polarized CBO (pCBO) Standard CBO converges to a single global optimum, but in many applications we may prefer multiple good local optima—for instance, when optimizing a surrogate objective (e.g., performance in a simulated environment) while caring about a different metric at deployment. To address this, Bungert et al. (2025) proposed Polarized CBO, which replaces the global consensus point with a local consensus point for each particle. The local consensus point for particle xit at time t is defined as
Optimization via Integration (OVI) (Andrieu et al., 2024) replaces the expectation in ES with a “log-sum-exp” aggregation, a relaxation of the minimum. While ES minimizes the expected loss Eπ [F (x)], OVI minimizes J OVI (θ) = − log Eϵ∼N (0,σ2 I) [exp(−βF (θ + ϵ))] , where β > 0 controls the sharpness of the minimum approximation. The OVI updates correspond to gradient descent on J OVI with step size σ 2 : PN i −βF (xi ) t xt e θt+1 = θt − σ 2 ∇θ J OVI (θt ) = Pi=1 , (OVI) i N −βF (xt ) i=1 e
PN mip,t =
i.i.d.
where xit ∼ N (θt , σ 2 I). This computes a weighted average of samples, with weights wti ∝ exp(−βF (xit )) favoring low-loss regions. OVI can also be derived from a Bayesian perspective (see Appendix A.3). In practice, Andrieu et al. (2024) set β = 1/σ̂F , where σ̂F is the sample standard deviation of fitness values.
j j j i j=1 xt exp(−βF (xt ))k(xt , xt ) . PN j j i j=1 exp(−βF (xt ))k(xt , xt )
(3)
The localizing kernel is typically Gaussian: k(x, y) = exp −∥x − y∥2 /(2κ2 ) . The bandwidth κ controls the effective interaction range between particles. This creates polarization: particles in different regions of the domain effectively form independent subpopulations, each converging to its own local optimum. For k = 1, we recover standard CBO. Particles are then updated via the (CBO) updates with mip,t replacing the global consensus mt for each particle xit .
3.2. Nonparametric Methods Consensus Based Optimization (CBO) (Pinnau et al., 2017; Carrillo et al., 2021) is a particle-based method for global optimization. N interacting particles x1t , . . . , xN t ∈ Rd , evolve according to the dynamics specified by the system of SDEs dxit = −λ(xit − mt )dt + σ∥xit − mt ∥dWti ,
i.i.d.
ϵit ∼ N (0, I). (CBO) The hyperparameter λ ∈ (0, 1] controls the strength of attraction to the consensus point, while σ controls exploration strength. A distinguishing feature of CBO compared to other prior particle-based methods is its amenability to the mean-field limit N → ∞, which allows establishing convergence guarantees to the global minimum under suitable assumptions (Pinnau et al., 2017; Carrillo et al., 2021).
(NES)
Clustered CBO (cCBO) While pCBO can find multiple optima in low dimensions, Bungert et al. (2025) show it struggles when d ≥ 10 due to the curse of dimensionality affecting kernel-based locality. To improve scalability, they propose cCBO, which maintains Nc cluster centers c1 , . . . , cNc and soft assignments pi,j t ∈ [0, 1] representing the probability that particle i belongs to cluster j. Each particle xit is attracted to its individual consensus point, defined as the weighted combination of cluster centers:
(1)
where Wti are independent Brownian motions and mt is the consensus point, defined as the weighted average PN i i i=1 xt exp(−βF (xt )) mt = P . (2) N i i=1 exp(−βF (xt ))
mic,t =
Nc X j pi,j t ct . j=1
The parameter β > 0 controls the sharpness of the weighting: as β → ∞, the consensus point converges to the particle with lowest objective value. The drift term attracts each
Particles are then updated via CBO dynamics (CBO) with mic,t replacing the global consensus mt . The cluster assign3
Bridging Spherical Black-Box Optimizers
ments and centers are updated as: rti,j k(xit , cjt ) , rti,j = pi,j t ← PNc i,j ′ j′ i j ′ =1 rt k(xt , ct ) PN i i,j xt pt exp(−βF (xit )) . cjt ← Pi=1 N i,j i i=1 pt exp(−βF (xt ))
pi,j t maxj ′ pi,j t
search distribution) they yield identical updates when using Spherical Gaussian distributions. Comparing (SS,NES), we see that both reduce to the same gradient estimator, differing only by a σ 2 scaling absorbed into the learning rate η. This equivalence justifies referring to both methods simply as ES throughout this work.
α ′
(4) The parameter α > 0 controls assignment sharpness: larger α encourages particles to commit to a single cluster. Each cluster center cjt evolves as a consensus point over its assigned particles, enabling simultaneous optimization toward Nc distinct optima.
4.2. CH ≈ CBO While CH uses state-independent noise σ ′ and CBO uses distance-scaled noise σ∥xit − mt ∥, these formulations become approximately equivalent in high dimensions. The radii ∥xit − mt ∥ concentrate across particles (and, in clustered variants, q P within each cluster) around a common value N 1 i i 2 rt = i=1 ∥xt − mt ∥ , so σ∥xt − mt ∥ ≈ σrt for N most particles. Consequently, CBO with λ = 1 behaves similarly to CH with noise scale σt′ = σrt . These heuristic arguments can be made precise by combining the mean-field limit N → ∞ (yielding a particle law ρt ) with the highdimensional limit d → ∞ (yielding the concentration of the empirical RMS radius rt toward a deterministic schedule).
Consensus Hopping (CH) Consensus Hopping (Riedl et al., 2024) is defined by the update xit+1 ← mt + σ ′ ϵit ,
ϵit ∼ N (0, I),
(CH)
where mt is defined as in (2). This resembles CBO with λ = 1, but uses a state-independent diffusion coefficient σ ′ rather than the distance-dependent term σ∥xit −mt ∥. In high dimensions, the radii ∥xit −mt ∥ concentrate across particles, making the two formulations approximately equivalent, as we discuss in Section 4.2.
4.3. DE ≈ pCBO We show that Diffusion Evolution is more accurately understood as pCBO with a time-dependent kernel. In particular, we note that (5) is not a valid approximation to the optimal denoiser Eπ [x0 | xt ] (see Appendix A.4 for a brief review of denoising diffusion models). As N → ∞, the MC approximation in (5) converges to
Diffusion Evolution (DE) Zhang et al. (2025) proposed Diffusion Evolution (DE), motivated by denoising diffusion models (Sohl-Dickstein et al., 2015; Song et al., 2021b). DE computes a local consensus for each particle using a time-dependent kernel: PN j j j i j=1 xt exp(−βF (xt ))kt (xt , xt ) i x̂0 = PN , (5) j j i j=1 exp(−βF (xt ))kt (xt , xt )
R x0 π(x0 )ρt (x0 )qαt (xt | x0 )dx0 x̂0 → R , π(x0 )ρt (x0 )qαt (xt | x0 )dx0 where ρt is the distribution of (infinitely many) particles at time t. Comparing with Eπ [x0 | xt ], there is an extra factor ρt . That is, the denoiser assumes a time-dependent data distribution proportional to πρt , instead of being proportional to π. A correct particle-based importance-sampled approximation would require weights 1/ρt , but these are unavailable since ρt itself is unknown. The denoised estimate x̂i0 in (5) is exactly a pCBO local consensus point (3) with kernel (6). The time-dependence of kt provides an annealing schedule: as αt → 1, the bandwidth (1 − αt ) → 0, transitioning from global interaction to local interaction.
where the kernel √ ∥x − αt y∥2 kt (x, y) = exp − 2(1 − αt )
(6)
derives from the diffusion forward process with schedule αt . Particles are updated using the DDIM sampler (Song et al., 2021a): √ q xi − αt x̂i0 √ xit−1 = αt−1 x̂i0 + 1 − αt−1 − σt2 · t√ +σt ϵit , 1 − αt (DE) i i.i.d. with ϵt ∼ N (0, I). See Appendix A.4 for the full derivation from the diffusion model perspective.
DE employs the DDIM sampler instead of the EulerMaruyama scheme typically employed in CBO. Both are valid discretization schemes. Moreover, DE employs a deterministic diffusion schedule, in contrast to (CBO). However, we already noted that in high-dimensional settings the statedependency of CBO’s diffusion coefficient is often limited. In summary, we argue that DE is most easily understood as a time-inhomogeneous implementation of pCBO, with different discretization scheme and noise schedules.
4. Connections 4.1. SS ≡ NES Although Stochastic Smoothing and Natural Evolution Strategies arise from different motivations (SS from constructing a differentiable surrogate, NES from optimizing a 4
Bridging Spherical Black-Box Optimizers = 0.1, JES
= 0.3, JES
= 0.4, JES
= 0.0
= 0.01
= 0.03
= 0.1
= 0.2
= 0.3
1.00
2 1 0
0.75 Y
Y
= 0.0, JES
0.50
2 0.25
Y
= 0.01, JOVI
= 0.1, JOVI
= 0.3, JOVI
= 0.4, JOVI
0.00
2 1 0
1.00
2
0.75
Y
= 0.1, J
OVIZ
= 0.3, J
OVIZ
= 0.4, J
OVIZ
Y
= 0.01, J
OVIZ
0.25
2 1 0
0.00 0.0
2 2.5 0.0 X
2.5
2.5 0.0 X
2.5
2.5 0.0 X
2.5
2.5 0.0 X
2.5
J
0.0
0.5 X
1.0
0.0
0.5 X
1.0
ES-OVI To allow us to control the convergence behavior of the optimizer depending on the application, we propose to interpolate the two methods via a convex combination of their gradients:
Both ES and OVI use a symmetric Gaussian proposal distribution N (θ, σ 2 I). Their difference lies in the surrogate objectives J ES and J OVI : OVI
1.0
as illustrated by the Rosenbrock example in Figure 2, an excessive preference for flat regions can be detrimental, causing convergence to suboptimal solutions.
4.4. ES ↔ OVI
ES
0.5 X
Figure 3. ES-OVI lets us control the flatness of the optimum. Markers indicate the minimum of J α on the Rosenbrock function for different values of α, yellow (α = 0, ES) to red (α = 1, OVI).
Figure 2. ES prefers flat basins, OVI sharp optima. Objectives of ES J ES , OVI J OVI (with β = 1), and fitness-scaled OVI (with β = 1/σ̂F ), for different noise levels σ on the Rosenbrock function. The true minimum is at (1, 1); the red x marks the minimum of each smoothed objective. Under J ES , the minimum moves toward the flat valley around [0, 0], while it stays closer to the true minimum under J OVI .
J
0.50
∇J α (θ) = α∇J OVI (θ) + (1 − α)∇J ES (θ),
(θ) = Eϵ∼N (0,σ2 I) [F (θ + ϵ)]
α ∈ [0, 1].
Since both objectives use the same samples ϵit and function evaluations F (θ + ϵit ), this combination requires no additional function evaluations. As shown in Figure 3, varying α smoothly interpolates between the convergence points of ES (α = 0) and OVI (α = 1), allowing us to obtain a solution with a desired flatness. We refer to this method as ES-OVI.
(θ) = − log Eϵ∼N (0,σ2 I) [exp(−βF (θ + ϵ))] .
At first glance, both objectives appear similar: each samples perturbations from a symmetric Gaussian and aggregates function values. When F is locally flat around θ, the two objectives indeed yield similar landscapes. However, they differ markedly near sharp minima: ES smooths them out, whereas OVI tends to preserve or even accentuate them.
4.5. OVI ≡ CH While nonparametric CBO and parametric OVI may appear quite dissimilar, we observe that Consensus Hopping is equivalent to OVI. Expanding the (CH) update with the definition of the consensus point yields PN j j j=1 xt exp(−βF (xt )) xit+1 = PN + σ ′ ϵit , j j=1 exp(−βF (xt ))
To illustrate, we visualize both surrogate objectives on the Rosenbrock function F (x, y) = (1 − x)2 + 100(y − x2 )2 (Rosenbrock, 1960), which has a sharp global minimum at (1, 1) and a broad, suboptimal valley near the origin. Figure 2 shows that as σ increases, the minimum of J ES shifts toward the origin where F (0, 0) = 1. In contrast, the minimum of J OVI with β = 1 remains close to the true optimum, even for large σ. When using adaptive scaling β = 1/σ̂F , OVI behaves more similarly to ES: Scaling by the inverse standard deviation attenuates sharp minima (which induce high variance in F ) while leaving flat regions relatively unchanged.
which is precisely the (OVI) update (see Appendix A.5 for a formal proof). This equivalence is significant: Particle-based methods typically struggle in higher-dimensional settings, as we will observe in Section 5.1, while OVI performs competitively with ES on higher-dimensional Brax tasks (Figure 6). On the other hand, both OVI and ES, being parametric methods with Gaussian distributions, are limited to finding a single optimum, whereas cCBO can discover multiple optima. We
Flat minima are often preferable to sharp ones, as they are associated with better generalization in both supervised learning (Hochreiter & Schmidhuber, 1997; Foret et al., 2021) and reinforcement learning (Lee & Yoon, 2025). However, 5
Bridging Spherical Black-Box Optimizers
1 0
Generation
3 2 1 AdaPol CBO (const) cCBO ES-OVI ES OVI pCBO Sep-CMA-ES
Final Max Fit.
×103
1
−0.1
3
0.8
−0.2
2
0.6
−0.3
500 1,000
0 ×103 1.4 1.2 1 0.8 0.6
500 1,000 Generation
Halfcheetah
×103
AdaPol CBO (const.) cCBO ES-OVI ES OVI pCBO Sep-CMA-ES
1 0
500 1,000 Generation
×103
0
Generation
×103 4
−0.1
500 1,000
3
−0.2
2
−0.3
1 AdaPol CBO (const.) cCBO ES-OVI ES OVI pCBO Sep-CMA-ES
2
Reacher
×103 0
AdaPol CBO (const.) cCBO ES-OVI ES OVI pCBO Sep-CMA-ES
Cumm. Max Fit.
3
Hopper
×103 1.2
AdaPol CBO (const.) cCBO ES-OVI ES OVI pCBO Sep-CMA-ES
Acrobot
×103
Figure 4. AdaPol improves upon CBO in higher dimensional tasks. OVI is competitive with CMA-ES. Evaluation on Brax tasks with 10 seeds each. The shaded areas show 95% CIs across 10 seeds, hyper-parameters were optimized for each method per each task.
Rastrigin
Sphere Min Fitness
thus combine both approaches to obtain a method capable of finding multiple optima in higher-dimensional problems. AdaPol & SchedPol Since CH is equivalent to OVI, we can interpolate between the two regimes by varying λ: We propose starting with λ = 1 (CH/OVI behavior) to quickly reach a promising region, and then reduce λ to enable multimodal exploration via cCBO. A fixed schedule for λ is an option, but determining when to transition is problemdependent. We thus propose an adaptive scheme inspired by JADE (Zhang & Sanderson, 2007). At each iteration, we allocate particles to two strategies, CH (λ = 1) and cCBO (λ < 1), in proportion to the number of successful particles each strategy produced over the previous NG generations, where success is defined as having fitness in the top p% of the current population. If either strategy’s success rate drops to zero, we allocate a small fraction of particles to it for continued exploration. We refer to this adaptive method as AdaPol and to the fixed-schedule variant as SchedPol.
103
10
10−8
10−6
−19
−14
101 10
Min Fitness
LinearSlope 102
2
10 AttractiveSector
Rosenbrock
106
105
10−3
10−4
10−12 0
10−13 1,000 2,000 0 Generation
1,000 2,000
ES CBO CBO (const) OVI DiffEvo CMA-ES Sep-CMA-ES pCBO ES-OVI
Generation
Figure 5. Hybrid optimizers can outperform base methods. Evaluation on a selection of 2D BBO tasks. Note that in multiple problems, ES-OVI performs better than either ES or OVI. The shaded areas show 95% CIs across 10 random seeds, hyperparameters were optimized for each method in each problem.
5.1. Benchmarks
5. Experiments
To our knowledge, neither CBO nor OVI has been evaluated on the popular BBO Benchmark (BBOB) (Hansen et al., 2009) or Brax (Freeman et al., 2021) benchmark. We thus evaluate both the existing and newly proposed methods in these two settings. We implement them in JAX (Bradbury et al., 2018), using the evosax library (Lange, 2022), reusing the existing implementations when available. We also evaluate a variant of CBO that uses a constant diffusion term σϵit instead of σ∥xit − mt ∥ϵit , denoted as “CBO (const.)”.
We first evaluate all methods on the common BBOB and Brax tasks, then investigate the behavior of ES-OVI specifically on the Brax tasks, and finally propose the application of multi-modal optimizers to LLM merging. Additional experimental details are shown in Appendix B, along with runtime comparisons.1 1 Our implementation is available at https://github. com/JohannesAck/bridging_optimizers.
6
Bridging Spherical Black-Box Optimizers
3.3
0.9 0.8 500 1,000
1 0.9 0.8
0 ×103
0
500 1,000 Generation
500 1,000 Generation
Humanoid
×103 2.5 2 1.5 1
ES (α = 0.0) α = 0.25 α = 0.5 α = 0.75 OVI (α = 1.0) 0
3.36
×103 2.8
×103
3.34
2.6
2
3.32
2.4
1.5
500 1,000 Generation
2.5
3.3
1 α α =0 = . α 0. 0 α = 025 = . α 0. 5 = 75 1. 0
Generation
α α =0 = . α 0. 0 α = 025 = . α 0. 5 = 75 1. 0
Final Max Fitness
2.4
3.2 0
×103
Halfcheetah
×103 2.8 2.6
3.25
α α =0 = . α 0. 0 α = 025 = . α 0. 5 = 75 1. 0
Max Fitness
1
Acrobot
×103 3.35
α α =0 = . α 0. 0 α = 025 = . α 0. 5 = 75 1. 0
Hopper
×103
Figure 6. ES-OVI allows us to tune the optimization behavior for each environment. Performance on Brax of ES-OVI with different interpolation-coefficients α. α = 1.0 corresponds to OVI, α = 0.0 corresponds to ES. 20 trials, 90% CIs.
MLP is the current state of the robot s and the output is the action a = f (x)(s) representing the torque applied to each actuator of the robot. We want to maximize the sum of the reward collected over an episode. The results are shown in Figure 4. We find that AdaPol tends to perform better than other CBO variants. Additionally, ES-OVI (α = 0.5) initially provides better performance than other methods, but tends to achieve a worse final result.
Black-Box Optimization Benchmark For lowdimensional problems, we evaluate each method on 23 of the BBOB benchmark problems (Hansen et al., 2009). The BBOB benchmark consists of a series of parametric problems, different in each initialization. As we use Brax for higher-dimensional evaluation, we here first use the 2D variants of the problems. For each problem, we optimize the hyperparameters of each method by grid search with 10 initializations each and select the best configuration based on mean fitness after 1000 generations. A subset of the problems is shown in Figure 5. In Appendix B.2, we show the results for the remaining tasks, along with the hyperparameter combinations considered for each method. We highlight some findings: In multimodal tasks, such as “RastriginOriginal”, the particle-based CBO and pCBO perform particularly well, as expected for particle-based methods. There are cases in which both OVI and ES fail, but ES-OVI performs well, such as “Attractive Sector”. Conversely, there are cases where it performs worse than either, such as “GriewankRosebrock”. This shows the utility of having another “knob” to control the convergence behavior, even when hyperparameters are optimized. We also repeated the experiments for higher dimensionalities and show the results in Appendix C.
5.2. ES-OVI To evaluate ES-OVI further, we perform experiments with different interpolation values α on the Brax tasks. The results are shown in Figure 6. While both ES (α = 0.0) and OVI (α = 1.0) have clear failure modes on different tasks, we find that setting an intermediate α = 0.75 performs well across tasks. We have also seen that on the Rosenbrock function, ES converges to a flatter region while OVI converges to the sharp minimum, with ES-OVI interpolating between them. To evaluate whether this behavior also occurs in higher dimensions, we perform experiments on the Acrobot task. After convergence, we evaluate the objective under Gaussian disturbances of the found parameters x, defined as E[F (x + ϵ)] with ϵ ∼ N (0, σ 2 I). The results are shown in Figure 7. OVI indeed finds a sharper, more sensitive solution than ES, while intermediate α values interpolate between them. As parameter sensitivity in reinforcement learning is connected to robustness to action disturbances or changed dynamics (Lee & Yoon, 2025), we also evaluate the performance under observation noise a = f (x)(s + ϵ) and action noise a = f (x)(s) + ϵ. The results, shown in Figure 7, indeed show a higher sensitivity to action noise of OVI than ES. Interestingly, ES-OVI achieves a better
Brax To evaluate the methods in a higher-dimensional setting, we use the Brax (Freeman et al., 2021) implementations of four classical control tasks. In these tasks, the goal is to find parameters x for a locomotion policy represented by a two-layer MLP with 32 units per layer. The problem dimensionality is thus approximately d ≈ 1000, with the exact dimensionality varying due to different input and output dimensionalities of each robot. The input of the 7
Bridging Spherical Black-Box Optimizers ×103
Parameter Robustness
×103
Scores for WizardMath vs Abel-7B
Observation Robustness
0.44
1.2
1.0 0.26
2.5
1
0. 05 0. 10 0. 30 0. 50 1. 00
Noise level σ
0. 1 0. 25 0. 4
00 1 0. 01
-0.0 0.01 0.02 0.04 0.05 0.06 Train Acc (64)
0.4
0.2
0.4
0.6
0.8
1.0
1.2
Weight WizardMath
2 0.
0.0
0.0 0.0
ES (α = 0.0) α = 0.25 α = 0.5 α = 0.75 OVI (α = 1.0)
2.5
0.09
0.6
0.2
Action Robustness
3
0.18
Noise level σ
MGSM-JA Accuracy
Mean Reward
×103
0.8
0. 1 0. 25 0. 4
2
0
Test Acc
2
Weight Abel-7B
3
0. 00 1 0. 01
Mean Reward
0.35
3
Best Cluster Worst Cluster
0.4
0.2
0.2
0
Noise level σ
0 SchedPol AdaPol
Figure 7. ES-OVI allows us to trade performance vs robustness. Robustness on the Acrobot task, trained with ES-OVI with different α values. Mean across 20 different random seeds with bootstrapped 95% CIs.
0.4
OVI
cCBO
SchedPol AdaPol
OVI
cCBO
Figure 8. Multimodal optimizers are advantageous in model merging with limited data. Top: Test and train accuracy for different merges, showing that model merging is a multi-modal problem. Bottom: Test accuracy when initializing centered on 0.0 (left) or with a broader init at 0.65 (right). Both show the mean across five random seeds, with bootstrapped 95% CIs. Multimodal cCBO is advantageous for the better init, our proposed SchedPol is competitive in both.
robustness than either OVI or ES under strong action or observation disturbances. This experiment also gives us some guidance on how we can approach picking the hyperparameter α: When we have less confidence in the fidelity of our function F , we may use a smaller α to obtain a more robust solution. When F is reliable, we can use a larger α to obtain a better performance.
optima on the small subset and evaluating each on the full training dataset only once at the end, we can significantly reduce the computational cost. We combine Mistral 7B (Jiang et al., 2023) with three finetuned versions, resulting in a 33-dimensional problem. Testing is done on 250 hand-translated MGSM questions (Shi et al., 2023), while we train on a fixed subset of 64 questions from the machine-translated GSM8K. To first show that model merging with a limited dataset is indeed a multimodal problem, we visualize the training and test accuracies for a 2D slice of the parameter space in Figure 8 (top). Next, we evaluate OVI, cCBO, AdaPol and SchedPol with a tuned λ schedule, each with two different initial distributions. In Figure 8 (bottom left), we show the test accuracy for two different initializations. The results show: When provided with a good initialization, treating model merging with limited data as a multi-modal problem and using cCBO performs significantly better than the unimodal OVI. Conversely, when initialized in a bad region of parameter space, cCBO performs worse than OVI. By combining both methods, either with a manually tuned SchedPol or AdaPol, we can perform competitively in both cases, as the OVI behavior first reaches a good region of parameter space in which cCBO can identify multiple optima. All methods are unable to match the accuracy of the method proposed by Akiba et al. (2025) due to overfitting the training dataset.
5.3. Evolutionary model merging To evaluate the utility of the combined parametric and particle-based AdaPol and SchedPol, we propose to use model merging with limited evaluations as a benchmark. Model merging (Wortsman et al., 2022; Ilharco et al., 2023) is a post-training method for LLMs, which combines multiple fine-tuned (FT) versions of the same model with different desirable characteristics. It interpolates P the parameters ϕ with weighting xi , ϕMerge = ϕBase + i xi (ϕFT,i − ϕbase ). These weights may be chosen by layer, resulting in a highdimensional problem, making the manual choice of x difficult. Evolutionary Model Merge (Akiba et al., 2025) thus uses CMA-ES to obtain a weighting. Their objective function evaluates the model on 1096 machine-translated Japanese samples from GSM8K, making each evaluation costly. Instead, we propose to evaluate each candidate only on a small, fixed subset of problems, significantly reducing the cost per candidate. Unfortunately, this turns the model merging problem into a multi-modal problem setting: Some optima of the objective overfit to the specific problems in the subset, while others generalize better. By obtaining multiple 8
Bridging Spherical Black-Box Optimizers
6. Related Work and Limitations
7. Conclusion
While we are not aware of any papers investigating the connection of ES, OVI, and CBO, other connections between optimizers have previously been investigated. Braun et al. (2025) investigated the combination of Stein Variational Gradient Descent (SVGD) with ES, replacing the gradient in SVGD with an estimate provided by ES or CMA-ES. Xu et al. (2019) proposed a combination of particle swarm optimization (PSO) and CMA-ES. However, PSO does not allow for any theoretical analysis of its convergence and the combination, while effective, appears heuristic. We thus focus on CBO and its variants instead, due to their advantageous theoretical properties and their relations to ES and OVI. Learned Evolution Strategy (LES) (Lange et al., 2023) investigated how the parameter update in ES methods can be meta-learned. In principle, LES could learn both the update rule of ES as well as OVI, and their combinations in our proposed ES-OVI method. However, by uncovering these connections manually we can explicitly control their properties. LES does not extend to particle based methods, but it would be interesting to consider similar meta-learning methods applied to the interaction matrix in our setup. Carrillo et al. (2021) introduced multiple changes for CBO to allow application to high-dimensional problem settings. However, it does not permit us to obtain multiple optima in high-dimensional problems, which aim for by combining cCBO and OVI. Information Geometric Optimization (IGO) (Ollivier et al., 2017) provides a unified framework for multiple black box optimizers, including NES, CMA-ES, the cross-entropy method, and related parametric methods. To our understanding, neither OVI nor the non-parametric CBO or DiffEvo, which we consider, fit into the IGO framework.
We investigated connections between multiple parametric and non-parametric Black-Box Optimizer. We proposed a unified view, revealing that the considered BBO methods differ primarily in two design choices: aggregation (expectation vs. log-sum-exp, controlling sharpness preference) and consensus scope (global vs. local, controlling modality). By interpolating along these axes, we obtained new methods that can outperform either of their endpoints. We hope that these perspectives will help guide both method selection for practitioners and future algorithm design.
Impact Statement Our work focuses on black-box optimizers. Optimizers are general purpose tools with many different applications, both positive and negative ones. There are thus many potential societal consequences, none of which we feel must be specifically highlighted here.
Acknowledgements We would like to thank Robert Lange, Stefania Druga, and Maxence Faldor for helpful advice and discussions.
References Akiba, T., Shing, M., Tang, Y., Sun, Q., and Ha, D. Evolutionary optimization of model merging recipes. Nature Machine Intelligence, 7:195–204, 2025. Andrieu, C., Chopin, N., Fincato, E., and Gerber, M. Gradient-free optimization via integration, 2024. URL https://arxiv.org/abs/2408.00888.
MERGE3 (Mencattini et al., 2025) proposes to improve the sample efficiency of Evolutionary Model Merging by using a reduced dataset for training, similar to our experiments on evolutionary model merging. Orthogonally to our work, they focus on using item response theory to more accurately predict the performance on the whole dataset from evaluations on the limited dataset, while we instead change the optimizer to allow us to obtain multiple optima.
augmxnt. shisa-gamma-7b-v1, 2023. https://huggingface.co/augmxnt/ shisa-gamma-7b-v1.
URL
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/jax-ml/jax.
6.1. Limitations A key limitation of our work is that we consider only optimizers based on spherical Gaussian distributions and thus cannot cover methods which adapt the covariance matrix during training, such as CMA-ES (Hansen, 2016). Our hybrid methods also introduce new hyper-parameters, NG in the case of AdaPol and α in the case of ES-OVI. While we found NG to not be very sensitive, α heavily affects the convergence characteristics of ES-OVI, as discussed in Section 5.2. We believe that addressing these limitations would be interesting avenue for future work.
Braun, C. V., Lange, R. T., and Toussaint, M. Stein Variational Evolution Strategies. In UAI. arXiv, 2025. URL https://proceedings.mlr.press/ v286/braun25a.html. Bungert, L., Roith, T., and Wacker, P. Polarized consensusbased dynamics for optimization and sampling. Mathematical Programming, 211:125–155, 2025. Carrillo, J. A., Jin, S., Li, L., and Zhu, Y. A consensus-based global optimization method for high dimensional machine 9
Bridging Spherical Black-Box Optimizers
learning problems. ESAIM: Control, Optimisation and Calculus of Variations, 27:S5, 2021.
Lange, R. T. evosax: Jax-based evolution strategies, 2022. URL https://arxiv.org/abs/2212.04180.
Chern, E., Zou, H., Li, X., Hu, J., Feng, K., Li, J., and Liu, P. Generative ai for math: Abel. https://github. com/GAIR-NLP/abel, 2023.
Lange, R. T., Schaul, T., Chen, Y., Zahavy, T., Dallibard, V., Lu, C., Singh, S., and Flennerhag, S. Discovering Evolution Strategies via Meta-Black-Box Optimization. In ICLR, 2023. URL https://openreview.net/ forum?id=mFDU0fP3EQH.
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-Aware Minimization for Efficiently Improving Generalization. In ICLR, 2021.
Lee, H. K. and Yoon, S. W. Flat Reward in Policy Parameter Space Implies Robust Reinforcement Learning. In ICLR, 2025. URL https://openreview.net/forum? id=4OaO3GjP7k.
Freeman, C. D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O. Brax – A Differentiable Physics Engine for Large Scale Rigid Body Simulation. In NeurIPS Datasets and Benchmarks Track, 2021.
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. In ICRL, 2025. URL https: //openreview.net/forum?id=mMPMHWOdOy.
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. The language model evaluation harness, 07 2024. URL https://zenodo.org/records/12608602.
Mencattini, T., Minut, A. R., Crisostomi, D., Santilli, A., and Rodolà, E. MERGE$ˆ3$: Efficient Evolutionary Merging on Consumer-grade GPUs. In ICML, 2025. URL https://proceedings.mlr.press/ v267/mencattini25a.html.
Glasserman, P. Monte Carlo Methods in Financial Engineering, volume 53 of Stochastic Modelling and Applied Probability. Springer, New York, 2004.
Metz, L., Freeman, C. D., Schoenholz, S. S., and Kachman, T. Gradients are Not All You Need, 2022. URL http: //arxiv.org/abs/2111.05803.
Glynn, P. W. Likelihood ratio gradient estimation for stochastic systems. Commun. ACM, 33(10):75–84, 1990.
Nesterov, Y. and Spokoiny, V. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17(2):527–566, 2017.
Hansen, N. The CMA Evolution Strategy: A Tutorial, 2016. URL https://arxiv.org/abs/1604.00772. Hansen, N., Finck, S., Ros, R., and Auger, A. RealParameter Black-Box Optimization Benchmarking 2009: Noiseless Functions Definitions. Research Report RR6829, INRIA, 2009. URL https://inria.hal. science/inria-00362633.
Ollivier, Y., Arnold, L., Auger, A., and Hansen, N. Information-Geometric Optimization Algorithms: A Unifying Picture via Invariance Principles. JMLR, 2017. URL https://jmlr.org/papers/v18/ 14-467.html. Pinnau, R., Totzeck, C., Tse, O., and Martin, S. A consensusbased model for global optimization and its mean-field limit. Mathematical Models and Methods in Applied Sciences, 27(1):183–204, 2017.
Hochreiter, S. and Schmidhuber, J. Flat Minima. Neural Computation, 9(1):1–42, 1997. Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing Models with Task Arithmetic. In ICLR, 2023.
Rechenberg, I. Evolutionsstrategie: Optimierung technischer Systeme nach Prinzipien der biologischen Evolution. Problemata, 15. Frommann-Holzboog, StuttgartBad Cannstatt, 1973. ISBN 978-3-7728-0373-4.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7B, 2023. URL https: //arxiv.org/abs/2310.06825.
Riedl, K., Klock, T., Geldhauser, C., and Fornasier, M. Gradient is All You Need?, 2023. URL https://arxiv. org/abs/2306.09778. Riedl, K., Klock, T., Geldhauser, C., and Fornasier, M. How Consensus-Based Optimization can be Interpreted as a Stochastic Relaxation of Gradient Descent. In Differentiable Almost Everything Workshop, ICML 2024, 2024.
Katkovnik, V. Ya. and Kulchitsky, O. Yu. Convergence of a Class of Random Search Algorithms. Automation and Remote Control, 33:1321–1326, 1972. 10
Bridging Spherical Black-Box Optimizers
Rosenbrock, H. An automatic method for finding the greatest or least value of a function. The computer journal, 3 (3):175–184, 1960.
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y. Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch. In ICML, 2024.
Rubinstein, R. Y. The score function approach for sensitivity analysis of computer simulation models. Mathematics and Computers in Simulation, 28(5):351–379, 1986.
Zhang, J. and Sanderson, A. C. JADE: Self-adaptive differential evolution with fast and reliable convergence performance. In 2007 IEEE Congress on Evolutionary Computation, 2007.
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. Evolution Strategies as a Scalable Alternative to Reinforcement Learning, 2017. URL https://arxiv. org/abs/1703.03864. Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., Das, D., and Wei, J. Language Models are Multilingual Chain-of-Thought Reasoners. In ICLR, 2023. Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In ICML, 2015. Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In ICLR, 2021a. Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-Based Generative Modeling through Stochastic Differential Equations. In ICLR, 2021b. Spall, J. C. Introduction to Stochastic Search and Optimization: Estimation, Simulation, and Control. Wiley, Hoboken, NJ, 2003. Wierstra, D., Schaul, T., Peters, J., and Schmidhuber, J. Natural Evolution Strategies. In 2008 IEEE Congress on Evolutionary Computation, pp. 3381–3387, 2008. Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., and Schmidhuber, J. Natural Evolution Strategies. Journal of Machine Learning Research, 15(27):949–980, 2014. Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In ICML, 2022. Xu, P., Luo, W., Lin, X., Qiao, Y., and Zhu, T. Hybrid of PSO and CMA-ES for Global Optimization. In IEEE Congress on Evolutionary Computation, 2019. Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M. TIES-Merging: Resolving Interference When Merging Models. In NeurIPS, 2023. 11
Zhang, Y., Hartl, B., Hazan, H., and Levin, M. Diffusion Models are Evolutionary Algorithms. In ICLR, 2025. URL https://openreview.net/forum? id=xVefsBbG2O.
Bridging Spherical Black-Box Optimizers
A. Appendix A.1. Fitness Shaping Rather than using raw fitness values directly, implementations of ES and other BBO methods typically apply a fitness shaping transformation before computing updates. A common choice is rank-based shaping (Wierstra et al., 2014; Salimans et al., 2017), which replaces fitness values with a function of their ranks: TR (xi ) =
rank(xi ) − 0.5, N
(7)
where rank(xi ) ∈ {1, . . . , N } denotes the rank of sample xi among the N samples in the current generation (rank(xi ) = j if F (xi ) is the j-th smallest fitness value). Using (7) makes the optimization invariant to any monotonic transformation of the fitness function. An alternative is standardization, which centers and normalizes the fitness values: TZ (xi ) =
F (xi ) − µ̂F . σ̂F
where µ̂F and σ̂F are the sample mean and standard deviation of the fitness values {F (xi )}N i=1 . Both transformations yield updates with controlled magnitude, reducing sensitivity to the scale of the objective and facilitating learning rate selection (Wierstra et al., 2008). A.2. Stochastic Smoothing Given a smoothing distribution with full support on Rd and differentiable density π, the smoothed objective is JπSS (θ) = Eϵ∼π [F (θ + ϵ)]. Using the likelihood-ratio estimator (Rubinstein, 1986; Glynn, 1990; Glasserman, 2004), the gradient with respect to θ is ∇θ JπSS (θ) = Eϵ∼π [F (θ + ϵ)∇ϵ − log π(ϵ)] . Crucially, JπSS (θ) is differentiable in θ even when F is not. For a Gaussian smoothing distribution π = N (0, σ 2 I), the score simplifies to ∇ϵ − log π(ϵ) = ϵ/σ 2 , yielding ∇θ J SS (θ) =
1 Eϵ∼N (0,I) [F (θ + σϵ)ϵ] . σ
The gradient descent updates θt+1 = θt − η∇θ J SS (θt ) are implemented via the Monte Carlo estimator in (SS). A.3. Optimization Via Integration The OVI updates can be derived from a Bayesian perspective. A parametric family of reference distributions Π is chosen. At each iteration, the current search distribution πt ∈ Π is tilted toward regions of low loss via π̃t+1 (x) ∝ exp(−βF (x))πt (x), then projected back onto Π by minimizing the KL divergence KL(π̃t+1 ||π) over π ∈ Π. If Π = {N (θ, σ 2 I)}θ∈Rd with fixed σ 2 , the KL projection reduces to computing the mean of the tilted distribution: θt+1 = Eπ̃t+1 [x] =
Eπt [xe−βF (x) ] . Eπt [e−βF (x) ]
Employing the self-normalized importance sampling MC estimator yields (OVI). The OVI updates also correspond to gradient descent on J OVI with step size σ 2 . To see this, rewrite (OVI) as θt+1 = PN P i j θt + i=1 wi ϵi , where wi = e−βF (x ) / j e−βF (x ) and ϵi = xi − θt . Note that J OVI (θ) = − log Z(θ) where Z(θ) = Eϵ [e−βF (θ+ϵ) ], ϵ ∼ N (0, σ 2 I). Differentiating, ∇θ J OVI (θ) = −∇θ Z(θ)/Z(θ), where ∇θ Z(θ) = and thus ∇θ J OVI (θ) = − σ12
h i 1 Eϵ ϵe−βF (θ+ϵ) , 2 σ
i i 2 OVI (θt ). i w ϵ and θt+1 = θt − σ ∇θ J
P
12
Bridging Spherical Black-Box Optimizers
A.4. Denoising Diffusion Models and Diffusion Evolution √ √ In diffusion models, one defines a forward noising process xt = αt x0 + 1 − αt ϵ with ϵ ∼ N (0, I) and αt ∈ (0, 1] √ decreasing over time. The conditional distribution of xt given x0 is qαt (xt | x0 ) = N (xt ; αt x0 , (1 − αt )I). Given a fixed target distribution π(x0 ) ∝ exp(−βF (x0 )) over clean samples, the optimal denoiser minimizing E[∥x̂0 (xt ) − x0 ∥2 ] is the conditional expectation R x0 π(x0 )qαt (xt | x0 )dx0 mt (xt ) = Eπ [x0 | xt ] = R . (8) π(x0 )qαt (xt | x0 )dx0 Instead of computing this integral exactly, Diffusion Evolution approximates it using the current particle population via (5), where the kernel kt (x, y) ∝ qαt (x | y). A.5. Equivalence of CH and OVI We show that CH and OVI are the same particle method. Given particles {xit }N i=1 , CH defines the (global) consensus PN i i i=1 xt exp(−βF (xt )) . mt = P N i i=1 exp(−βF (xt )) CH then resamples xit+1 = mt + σϵit+1 ,
i.i.d.
ϵit+1 ∼ N (0, I).
OVI samples xit = θt + σϵit and updates the mean by the same weighted average: PN i i i=1 xt exp(−βF (xt )) θt+1 = P = mt , N i i=1 exp(−βF (xt )) then resamples xit+1 = θt+1 + σϵit+1 . Assume both methods start from the same θ0 and use the same Gaussian perturbations {ϵit }. Then at t = 0 they generate the same particles {xi0 }. If {xit } coincide at some iteration t, then both compute the same mt , and with the same perturbations they generate the same {xit+1 }. By induction, CH and OVI produce identical particle trajectories for all t, and their centers satisfy θt+1 = mt . A.6. ES as (MU) update We show how ES can be formulated as a particle update of the form (MU), provided that the sampling perturbations have zero empirical mean. Let θt ∈ Rd be the search distribution mean and σ > 0 fixed. Sample perturbations ϵit and candidates xit = θt + σϵit . PN Assume that the perturbations satisfy i=1 ϵit = 0, e.g. antithetic sampling with even N , for each t. The (mean-only) ES update corresponding to (NES) is N η X i i θt+1 = θt − gϵ, (9) N σ i=1 t t PN where gti = F (xit ). Under i=1 ϵit = 0, subtracting any constant baseline from the scores leaves the ES update unchanged, so (9) is equivalent to the baseline-centered form N
θt+1 = θt − Note that θt can be recovered from {xit }: N1
η X i (g − ḡt )ϵit , N σ i=1 t
N
ḡt =
1 X i g. N i=1 t
PN
i i=1 xt = θt .
Define per-particle weights wti =
η 1 − (g i − ḡt ), N N σ2 t 13
(10)
Bridging Spherical Black-Box Optimizers
which satisfy
PN
i i=1 wt = 1. For the consensus mt =
mt = θt + σ
N X
PN
i i i i i=1 wt xt , using xt = θt + σϵt and
i i ϵt = 0, we obtain
P
N
η X i (g − ḡt )ϵit = θt+1 . N σ i=1 t
wti ϵit = θt −
i=1
Therefore ES can be written as a particle-only update: (i) compute mt from particles using weights (10), (ii) resample xit+1 = mt + σϵit+1 . This matches (MU) with Kti,j = 1, µt = 0, λt = 1, and s(∆) = 1. P i.i.d. If ϵit ∼ N (0, I) without enforcing i ϵit = 0, then the particle mean is N
N 1 X i 1 ϵ̄t = ϵ ∼ N 0, I . N i=1 t N
1 X i θ̂t = x = θt + σϵ̄t , N i=1 t
If one replaces θt by θ̂t , as in the particle ES update from (MU), the resulting iterates differ from the true ES updates from (9) by terms proportional to ϵ̄t . So the particle representation for ES remains a good approximation even without antithetic sampling, as long as N ≫ 1. A.7. cCBO as (MU) update cCBO (Bungert et al., 2025) maintains soft assignments pi,c ∈ [0, 1] of particle i to cluster c ∈ {1, . . . , NC }, with t PNC i,c p = 1, and defines cluster centers as fitness-weighted averages of the particles. We show that its per-particle c=1 t consensus can be written in the MU form (MU) by choosing a cluster-induced low-rank interaction matrix Kt . Let Ψ(F ) = exp{−βF } and define the cluster normalizers Ztc :=
N X j pj,c t Ψ(F (xt )). j=1
Define the cluster-induced interaction (for indices i, j) Kti,j :=
NC i,c j,c X p p t
c=1
t
Ztc
.
(11)
Then the MU unnormalized weights satisfy j i,j ai,j t = Ψ(F (xt ))Kt =
NC j X pj,c t Ψ(F (xt )) pi,c , t Ztc c=1
where NC N X X ai,j pi,c t = t = 1, j=1
c=1
hence wti,j = ai,j t . The corresponding MU consensus becomes NC N j,c j X X p Ψ(F (x )) t , mit = wti,j xjt = pi,c xjt t t c Z t c=1 j=1 j=1 | {z } N X
=: cct
which equals the cCBO consensus mic,t = assigned to cluster c.
PNC
i,c c c c=1 pt ct , where each center ct is a fitness-weighted average over the particles
cCBO additionally specifies how to update the assignments pi,c t over time; in our framework this corresponds to specifying the time evolution of the interaction matrix Kt . 14
Bridging Spherical Black-Box Optimizers Table 2. Hyperparameters used in BBOB tasks. Other hyper-parameter were optimized by grid-search over values shown below.
Hyperparameter
Value
Generations Population DE Mapping DE α schedule ES-OVI α
2000 256 Energy cosine 0.5
B. Experiment Details B.1. Implementation Details We implement all methods in Jax (Bradbury et al., 2018) based on the existing methods in evosax (Lange, 2022). Due to Jax’s just-in-time compilation, this allows the evaluation of multiple candidates x and even multiple algorithm hyperparameters in parallel, significantly decreasing the computational cost of hyper-parameter optimization. Unless stated otherwise below, we use2the default hyper-parameters provided by evosax. For pCBO and cCBO we use Gaussian kernels ∥x−y∥ k(x, y) = exp − 2κ2 2 , with the squared Euclidean norm ∥ · ∥22 . AdaPol, SchedPol In AdaPol we follow JADE (Zhang & Sanderson, 2007) and allocate particles to each of two strategy (λ = 1 and λ = 0.1) proportionally to the success rate of each over the last NG generations. The success rate is defined as the percentage of candidates belonging to each strategy in the top p% of the total particles. To reallocate particles to strategies as needed we change as few particles as possible, and when changing from OVI behavior to cCBO behavior randomly initialize the cluster assignments pi,j . When all particles are assigned to OVI, we randomly assign 1/3 of particles to cCBO for exploration. In this case, we randomly sample new cluster centers cj uniformly from the OVI particles reassigned to OVI and randomly initialize all cluster assignments pi,j , by sampling uniform pi,j ∼ U(0, 1) and then normalizing for each particle. When using SchedPol this only happens once when changing from λ = 1 to λ < 1. We also use this initialization in the first step of our cCBO implementation. B.2. BBOB We show the full results on all tested BBOB functions in Figure 9. B.2.1. H YPERPARAMETER O PTIMIZATION Fixed hyper-parameters are shown in Table 2. We tune hyper-parmeters for each problem with 10 initializations and grid-search over the following parameters. • ES: σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, lr ∈ {0.00001, 0.00003, 0.0001, 0.0003, 0.001, 0.003} • CMA-ES: σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, cM ∈ {0.8, 0.9, 1.0, 1.1, 1.2} • Sep-CMA-ES: σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, cM ∈ {0.8, 0.9, 1.0, 1.1, 1.2} • Diffusion Evolution σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, fitnessmap temperature ∈ {0.6, 0.7, 0.8, 0.9, 1.0} • CBO σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, λ ∈ {0.001, 0.1, 0.2, 0.3, 0.5} • pCBO σ ∈ {0.0001, 0.001, 0.01}, λ ∈ {0.1, 0.2, 0.001, 0.3, 0.5}, κ ∈ {1, 2, 5} • OVI σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, β ∈ {0.5, 0.75, 1.0, 1.25, 1.5} • ES-OVI σ ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01}, lr ∈ {0.00001, 0.00003, 0.0001, 0.0003, 0.001, 0.003} • AdaPol σ ∈ {0.001, 0.003, 0.01, 0.03, 0.1, 0.3} In addition, for all methods we also evaluate using the raw scores vs rank-based fitness transformation, and constant σ vs decayed σ. We choose the best hyper-parameters by best mean fitness after 1000 generations. 15
Bridging Spherical Black-Box Optimizers Table 3. Hyperparameters used in Brax tasks. Other hyper-parameter were optimized by grid-search over values shown below.
Hyperparameter
Value
Generations Episode Length Population ES-OVI α pCBO, cCBO λ cCBO NC
1000 500 256 0.5 0.5 4
B.3. Brax experiments We run each environment for 500 steps per episode, diverging from the commonly used 1000 steps per episode to reduce computational cost. We were unable to obtain a satisfying result using Diffusion Evolution on the Brax tasks and thus excluded it from the results. For all CBO based methods we here use the constant noise term as in CBO (const.), as we found it to be beneficial to prevent premature convergence to suboptimal policies. B.3.1. H YPERPARAMETER OPTIMIZATION We optimize the hyper-parameters for each approach in each task by grid-search, other hyper-parameters are shown in Table 3. For the Brax tasks, we use 10 seeds each for Acrobot, Hopper and Reacher and 5 seeds each for Half-Cheetah. We choose the best hyper-parameter by max-fitness after 1000 generations. We perform grid-search over the following hyperparameters: • ES: σ ∈ {0.01, 0.03, 0.1, 0.3}, lr ∈ {0.0001, 0.0003, 0.001, 0.003, 0.01, 0.03, 0.1} • Sep-CMA-ES: σ ∈ {0.1, 0.3, 1.0} • CBO σ ∈ {0.01, 0.03, 0.1, 0.3}, λ ∈ {0.1, 0.2, 0.001, 0.3, 0.5} • pCBO σ ∈ {0.01, 0.03, 0.1, 0.3}, κ ∈ {0.1, 1, 10, 100} • cCBO σ ∈ {0.01, 0.03, 0.1, 0.3}, κ ∈ {0.1, 1, 10, 100} • OVI σ ∈ {0.001, 0.003, 0.01, 0.03, 0.1, 0.3, 1.0} • ES-OVI σ ∈ {0.01, 0.03, 0.1, 0.3, 1.0}, lr ∈ {0.00005, 0.0001, 0.0003, 0.001, 0.003, 0.01} • AdaPol σ ∈ {0.001, 0.003, 0.01, 0.03, 0.1, 0.3} For ES we also tried to use a linear decay for the sampling distributions σ, as well as the learning rate, but we did not find a benefit of doing so and thus present the results with constant σ and learning rate. For AdaPol we use NG = 100, λ = 0.3. For OVI we use J OVIZ in the Brax experiments, i.e. β = 1/σˆF in each iteration. With a fixed β we were unable to obtain useful policies. B.4. Model Merging Details Following (Akiba et al., 2025), we implement model merging based on the DARE-TIES method and use the same base model Mistral-7B-v0.1 (Jiang et al., 2023) and three finetunes: WizardMath-7B-V1.1 (Luo et al., 2025), shisa-gamma-7b-v1 (augmxnt, 2023), and Abel-7B-002 (Chern et al., 2023). By implementing both the merging and evaluation on multiple GPUs, we can reduce the costly CPU-GPU transfers needed by other methods. However, this requires holding four 7B models in GPU memory, requiring 56GB for just the bf16 weights. To reduce this memory burden but avoid CPU-GPU transfers, we shard the base and finetuned models across 4GPUs and all-reduce the result across GPUs. We then perform the evaluation using a modified, stateful version of lmeval (Gao et al., 2024). For the black-box optimizers we reuse our Jax implementation based on evosax (Lange, 2022). Hyperparameters are shown in Table 4. We do not decay the noise during training as we found it not to be beneficial. Further, while we evaluated the combination of DARE (Yu et al., 2024) and 16
Bridging Spherical Black-Box Optimizers Table 4. Hyperparameters used for model merging experiments
Hyper-parameter
Values
Train Dataset Size Layer Reduction Factor Population Size Init range Weight range σ-init Generations Cluster NC cCBO κ cCBO α cCBO λ SchedPol OVI generations
64 4 128 [-0.1, 0.1] or [0.5,0.8] [-0.2,2.0] 0.4 60 4 6 4 0.3 40
Table 5. Runtime of different methods for 1000 iterations with population size 256, averaged across 10 trials each.
OpenAI-ES
OVI
ESOVI
CMA-ES
Sep-CMA-ES
CBO
DiffEvo
AdaPol
D=5 D=50 D=1000 D=5000
4.0s 4.0s 4.1s 4.2s
4.1s 4.1s 4.2s 4.3s
5.8s 5.8s 6.1s 6.1s
9.9s 10.3s 14.5s 149s
7.7s 7.7s 7.8s 8.0s
5.9s 6.0s 6.2s 6.5s
4.9s 5.0s 5.1s 5.7s
13.8s 13.8s 13.8s 14.0s
TIES (Yadav et al., 2023), we found in our experiments that the dropout introduced by DARE was not advantageous and thus only use TIES. Further, we do not use the top ∆ trimming used in TIES and only retain the sign consensus mechanism, as we did not find it to be beneficial in our setting. B.5. Runtime Comparison To investigate the impact of our proposed methods on the runtime of the algorithm, we run each method for 1000 iterations with a random fitness function and population size 256. We use a just-in-time compiled implementation in JAX and perform our benchmarks on an NVIDIA RTX 2080 Super GPU. The results are shown in Table 5. ESOVI takes significantly longer than either OpenAI-ES or OVI, and AdaPol takes significantly longer than OVI or CBO. However, in black-box optimization the evaluation of the function F we are optimizing is typically significantly more expensive than the black-box optimizer. For example, in our model merging experiments, evaluating a single candidate x takes 12 seconds on an NVIDIA H100 GPU, while the time taken to generate the new candidate population with 128 candidates is less than a second.
17
Bridging Spherical Black-Box Optimizers
C. BBOB results for different dimensionalities To investigate the performance of the different optimizers as we increase the problem dimensionality, we repeat the BBOB experiments on higher dimensional variants of the benchmark problems. We perform the same hyper-parameter optimization separately for each dimensionality, as discussed in Section B.2.1. As shown in Figures 10, 11, 12, 13, 14, we largely observe a similar trend for higher dimensional variants of the tasks.
18
Bridging Spherical Black-Box Optimizers
34
10
42
AttractiveSector
10
4
10
7
10
10
10
13
10
2
10
1
10
4
10
7
10
10
Discus
18
10
26
10
34
10
42
10
4 11
10
18
10
25
10
32
10
10
2
10
5
10
8
10
11
10
14
39
10 0 10
1
10
2
10
3
10
4
1
1000 1500 Generation
10
9
10
12
10
15
10 0 10
2
10
4
10
6
10 2 10 0 10
10
2
4
2000
10 0
2
10
4
10
6
10
8
10
10
0
500
1000 1500 Generation
1
10
4
10
7
10
10
10
13
10
16
10 1 10
1
10
3
10
5
10
7
10
9
2000
10
2
10
4
10
6
10
8
10
10
10
1
10
3
10
5
OpenAI-ES CBO CBO (Const.)
OVI DiffEvo CMA-ES t Sep-CMA-ES pCBO ES-OVI
0
500
1000 1500 Generation
2000
Figure 9. Results on BBOB tasks in 2D
19
10 10
10
10
18
10
26
10
34
10
42
RastriginRotated 10 1 10
2
10
5
10
8
10
11
GriewankRosenbrock 10 1
Gallagher21Hi 10 2
10 0 10
10
2
DifferentPowers
10 2
10 2 Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
10 0
500
6
Gallagher101Me
10 1
0
10
EllipsoidalRotated 10 6
10 2
SchaffersF7IllConditioned
10 1
Lunacek
10
3
10 1
RosenbrockRotated
10 0 10
SchaffersF7
10 2 Cummulative Min Fitness
Cummulative Min Fitness
Weierstrass 10 1
10 0
SharpRidge
10 3
10
10 1
10 5
10 3
Cummulative Min Fitness
10
11
Cummulative Min Fitness
2 10
Cummulative Min Fitness
Cummulative Min Fitness
10
10
BentCigar
10 6
10
10
9
10 2
RosenbrockOriginal
Cummulative Min Fitness
2 1
7
10 5 Cummulative Min Fitness
Cummulative Min Fitness
10
10
StepEllipsoidal
10 5
10
10
Cummulative Min Fitness
42
26
10
5
Cummulative Min Fitness
10
10
3
LinearSlope 10 2
Cummulative Min Fitness
35
1
10
Cummulative Min Fitness
28
10
10
18
10 10
Cummulative Min Fitness
10
2 10
BuecheRastrigin 10 3
1
Cummulative Min Fitness
10
21
10 10
Cummulative Min Fitness
7 14
Cummulative Min Fitness
Cummulative Min Fitness
10 10
RastriginOriginal
10 3
Schwefel
Cummulative Min Fitness
EllipsoidalOriginal 10 6
Cummulative Min Fitness
Sphere 10 0
10 4 10 2 10 0 10
2
10
4
10
6
Bridging Spherical Black-Box Optimizers
10 23 10 30 10 37 10 44
103 10 2 10 7 10
BuecheRastrigin
LinearSlope 2 × 102
12
10 17 10 22
AttractiveSector
100 10
2
10 4 10 6 10 8
StepEllipsoidal
Cummulative Min Fitness
10 9 10 16
RastriginOriginal 102 Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
EllipsoidalOriginal
108
Cummulative Min Fitness
Sphere 10 2
104
103
102
101
RosenbrockOriginal
102
RosenbrockRotated
EllipsoidalRotated
106
10 9 10 12
10 15
10 1 10 8 10 15 10 22 10 29 10 36
Weierstrass
10 3 10 8 10 13 10 18 10 23
100
10 2
10 4
10 1 10 3 10 5 10 7
SchaffersF7IllConditioned
RastriginRotated 100
101
10 7 10 14 10 21 10 28 10 35 10 42
GriewankRosenbrock
Schwefel 105
100 1
10 2 10 3 10 4
101
100
10 1
101 100 10 1 10 2 10 3 10 4
Lunacek
Cummulative Min Fitness
Cummulative Min Fitness
101
3
7
10 9 10 11
0
500
1000 1500 Generation
2000
10 5
Gallagher21Hi
10 5 10
10 3
101
10 1 10
10 1
OpenAI-ES CBO CBO (Const.)
10 1 10
3
OVI DiffEvo CMA-ES t Sep-CMA-ES
10 5 10
7
pCBO
10 9
ES-OVI
10 11 0
500
1000 1500 Generation
2000
10
4
103 102 101 100 10 1
Gallagher101Me 101
2
101
Cummulative Min Fitness
101
Cummulative Min Fitness
102 Cummulative Min Fitness
Cummulative Min Fitness
10
10
2
DifferentPowers
2
SchaffersF7
102
Cummulative Min Fitness
10 11
Cummulative Min Fitness
10 11
10
10 8
SharpRidge Cummulative Min Fitness
3
10 7
10
10 5
10 14
BentCigar
101 10
10 9 10 12
101 10 2
106
105
Cummulative Min Fitness
Cummulative Min Fitness
Discus
10 6
Cummulative Min Fitness
10 9 10 12
10 6
100 10 3
Cummulative Min Fitness
10 6
100 10 3
Cummulative Min Fitness
10
3
107
104
103
Cummulative Min Fitness
100
103
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
103
0
500
1000 1500 Generation
2000
Figure 10. Results on BBOB tasks in 5D
20
Bridging Spherical Black-Box Optimizers
EllipsoidalOriginal
RastriginOriginal
10 37 10 44
10 6 10 10 10 14
101 10 1 10 3 10 5 10 7
Cummulative Min Fitness
10 30
102 10 2
LinearSlope
105 Cummulative Min Fitness
10 23
Cummulative Min Fitness
10 9 10 16
BuecheRastrigin
103
106 Cummulative Min Fitness
Cummulative Min Fitness
Sphere 10 2
104 10
3
102 101
3 × 102
2 × 102
100
10 10
10 6 10 9 10 12
Discus
10 5 10 9 10 13
10 9 10 12
10 3 10 10 10 17 10 24 10 31 10 38
Weierstrass
10 5 10 8 10 11 10 14
10
0
10 2
10 4
10 6
101
100
10 3 10 5
101 100 10 1 10 2 10 3
100
10 2 10 3 10 4
102
0
500
1000 1500 Generation
2000
5
Gallagher21Hi Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
Gallagher101Me 101 10 1 10 3 10 5 10 7 10 9 0
500
1000 1500 Generation
2000
RastriginRotated
10 7 10 14 10 21 10 28 10 35 10 42
Schwefel 105
10 1
10
Lunacek
10 16
GriewankRosenbrock 101
10 4
10 1
10 8 10 12
102 Cummulative Min Fitness
10 4
10 1
SchaffersF7IllConditioned Cummulative Min Fitness
Cummulative Min Fitness
100 10 2
100 10 4
100
101
10 7
102
101
OpenAI-ES CBO CBO (Const.)
10 1
OVI DiffEvo CMA-ES t Sep-CMA-ES
10 3 10 5 10 7
pCBO ES-OVI
10 9 0
500
1000 1500 Generation
2000
Figure 11. Results on BBOB tasks in 7D
21
EllipsoidalRotated
104
DifferentPowers
102
SchaffersF7
102 Cummulative Min Fitness
101 10 2
SharpRidge
104
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
103 10 1
10 6
BentCigar
1011
107
100 10 3
108 Cummulative Min Fitness
10 7
10
3
104
Cummulative Min Fitness
10 4
100
RosenbrockRotated
103
Cummulative Min Fitness
102 10 1
10
103
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
105
RosenbrockOriginal
6
Cummulative Min Fitness
StepEllipsoidal
Cummulative Min Fitness
AttractiveSector
104 103 102 101 100
Bridging Spherical Black-Box Optimizers
EllipsoidalOriginal
10
21
10 28 10
35
10 42
10
3
BuecheRastrigin
10 1 10 5 10 9 10 13
10
LinearSlope
105 Cummulative Min Fitness
10 14
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
10
RastriginOriginal 103
107
7
1
10 1 10 3 10 5 10 7
Cummulative Min Fitness
Sphere 100
104 103 102 101
5 × 102
4 × 102
3 × 102
102 10 1 10 4 10 7 10 10
106
103
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
105
100 10 3 10 6 10 9 10 12
Discus
RosenbrockOriginal
106
103 100 10
3
10 6 10 9 10 12
BentCigar
RosenbrockRotated
100 10
3
10 6 10 9 10 12
SharpRidge
EllipsoidalRotated 106
103
Cummulative Min Fitness
StepEllipsoidal
Cummulative Min Fitness
AttractiveSector
102 10 2 10 6 10 10 10 14
DifferentPowers
RastriginRotated
1010
10 6 10 9 10 12
10 14 10 22 10 30 10 38
100
10 2
10 4
Weierstrass
SchaffersF7
10 6
10 5
101
100
Lunacek
101 100 10 1 10 2 10 3
Gallagher101Me
10
100 10 1 10 2
102
102 Cummulative Min Fitness
2 × 10
2
100 10
2
10 4 10 6 10 8
OpenAI-ES
100 10
CBO CBO (Const.)
2
OVI DiffEvo CMA-ES t Sep-CMA-ES
10 4 10 6
pCBO ES-OVI
10 8
6 × 101 0
500
1000 1500 Generation
2000
0
500
1000 1500 Generation
2000
10 28 10 35 10 42
Schwefel 105
Gallagher21Hi
102 Cummulative Min Fitness
Cummulative Min Fitness
3 × 10
2
10 21
GriewankRosenbrock
1
10 3
4 × 102
10 7 10 14
102 Cummulative Min Fitness
10 4
Cummulative Min Fitness
Cummulative Min Fitness
100
10 3
SchaffersF7IllConditioned 102
10 2
10 1
10 7
102 Cummulative Min Fitness
102
Cummulative Min Fitness
10 3
10
10 6
100
101
Cummulative Min Fitness
100
Cummulative Min Fitness
103
2
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
106
0
500
1000 1500 Generation
2000
Figure 12. Results on BBOB tasks in 10D
22
104 103 102 101 100
Bridging Spherical Black-Box Optimizers
Sphere
EllipsoidalOriginal
10 29 10 36 10 43
102 10 2 10 6 10 10 10 14
101 10 1 10 3 10 5
105
LinearSlope Cummulative Min Fitness
10 22
BuecheRastrigin Cummulative Min Fitness
10 8 10 15
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
RastriginOriginal 103
106
10 1
104 103 102
9 × 10
2
8 × 102
7 × 102
101
StepEllipsoidal
RosenbrockOriginal
10 3 10 6 10 9
100 10 3 10 6 10 9 10 12
Discus
100
10 6 10 9 10 12
10 6 10
9
10 12
100 10 4 10 8 10 12 10 16
DifferentPowers
RastriginRotated
10 12
10 25 10 33 10 41
SchaffersF7IllConditioned
10 7 10 11 10 15 10 19 10 23
2 × 102
2
2000
Cummulative Min Fitness
Cummulative Min Fitness
2
10 1 10 3 10 5 10 7
0
500
1000 1500 Generation
2000
10 21 10 28 10 35 10 42
GriewankRosenbrock
101 100 10 1 10 2
101
100
10 1
Schwefel
104 103 102 101 100
Gallagher21Hi
101
10 7 10 14
105
Gallagher101Me
4 × 102
10 5
10 3
Lunacek 6 × 102
10 3
102 Cummulative Min Fitness
5
10
3
10
10 1
10 7
102 Cummulative Min Fitness
Cummulative Min Fitness
10 3
1000 1500 Generation
10 4
SchaffersF7
10 1
500
10 2
10 6
101 101
0
100
Cummulative Min Fitness
10 9
10 17
10
100
1
Cummulative Min Fitness
10 6
10 9
2
Cummulative Min Fitness
10 3
Cummulative Min Fitness
100 10 3
SharpRidge
10 1
Weierstrass
Cummulative Min Fitness
104
10 3
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
100
10
108
103
BentCigar
103
3 × 10
EllipsoidalRotated
106
103
107
106
10
RosenbrockRotated
106
Cummulative Min Fitness
100
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
103
103
Cummulative Min Fitness
AttractiveSector 106
101
OpenAI-ES CBO CBO (Const.)
10 1
10 5
OVI DiffEvo CMA-ES t Sep-CMA-ES
10 7
ES-OVI
10 3
pCBO
0
500
1000 1500 Generation
2000
Figure 13. Results on BBOB tasks in 15D
23
Bridging Spherical Black-Box Optimizers
10 15 10 20
AttractiveSector
StepEllipsoidal
10 3 10 6 10 9
100 10
3
10 6 10 9 10 12
Discus
10
3
10 6 10 9 10 12
BentCigar
10 7 10 10 10 13
10 13 10 21 10 29 10 37
10 5
10 10
101
10 3 10 5 10 7
10
1
10 1 10 3 10 5 10 7
SchaffersF7IllConditioned
10 1
10 3
10 5
Lunacek
101 100 10 1 10 2
101
100
10 1
Gallagher21Hi 102
2 × 102
0
500
1000 1500 Generation
2000
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
102
3 × 102
10 20
RastriginRotated
100 10 2 10 4 10 6
0
500
1000 1500 Generation
2000
OpenAI-ES
100
CBO CBO (Const.)
10 4
OVI DiffEvo CMA-ES t Sep-CMA-ES
10 6
ES-OVI
10
102
100
10 2
10 4
Schwefel 105 104 103 102 101 100
Gallagher101Me
4 × 102
10 15
GriewankRosenbrock
10 3
6 × 102
10 10
102
102
101
Cummulative Min Fitness
Cummulative Min Fitness
10 3
10 7
100 10 5
DifferentPowers
10 1
SchaffersF7
10 1
10 4
SharpRidge
10 5
Weierstrass 101
102 10 1
Cummulative Min Fitness
10 4
10
EllipsoidalRotated 105
103
3
Cummulative Min Fitness
Cummulative Min Fitness
102 10 1
1.1 × 103
RosenbrockRotated
100
1011
105
1.2 × 103
105
103
Cummulative Min Fitness
100
103
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
103
102
RosenbrockOriginal 106
108 Cummulative Min Fitness
10 4
103
1.3 × 103
101
106
Cummulative Min Fitness
10 2
104
Cummulative Min Fitness
10 40
10 10
LinearSlope 1.4 × 103
Cummulative Min Fitness
10 33
100
105
Cummulative Min Fitness
10 26
100 10 5
BuecheRastrigin
Cummulative Min Fitness
10
19
Cummulative Min Fitness
Cummulative Min Fitness
Cummulative Min Fitness
10 12
RastriginOriginal 102
Cummulative Min Fitness
EllipsoidalOriginal 105
Cummulative Min Fitness
Sphere 102 10 5
2
pCBO
0
500
1000 1500 Generation
2000
Figure 14. Results on BBOB tasks in 20D
24