arXiv:2609.06830v1 [cs.LG] 6 Sep 2026
Constrained Bayesian Optimization for Hierarchical Federated Learning in IoT Networks for Plant Disease Classification Athanasios Papanikolaou‡ , Athanasios Tziouvaras† , Apostolos Xenakis† , Periklis Chatzimisios∥ , Shameem A. Puthiya Parambath+ , George Floros§ , Enrica Zereik* , Ivan Petrovic‡ , and Fabio Bonsignorio‡ ‡ University of Zagreb, Croatia, emails: {athanasios.papanikolaou, ivan.petrovic, fabio.bonsignorio}@fer.unizg.hr † University of Thessaly, Greece, emails: {attziouv, axenakis}@uth.gr ∥ International Hellenic University, email: [email protected] + University of Glasgow, UK, email: [email protected] § Trinity College Dublin, Ireland, email: [email protected] * Italian National Research Council, Institute of Marine Engineering, Italy, email: [email protected] Abstract—The deployment of Hierarchical Federated Learning (HFL) in resource-constrained Internet of Things (IoT) environments requires careful configuration to balance predictive performance with energy consumption and execution time. This challenge is particularly relevant to smart agriculture, where distributed IoT devices can support automated plant disease classification while operating under limited computational and communication resources. This paper presents a constrained Bayesian Optimization framework for the efficient configuration of HFL deployments. The proposed approach jointly explores the deep learning backbone architecture, aggregation strategy, and number of communication rounds, while the federation size is determined according to the spatial coverage requirements of the agricultural deployment. A weighted objective function captures user-defined trade-offs among energy consumption, execution time, and predictive performance, while explicit constraints ensure compliance with deployment-specific resource and accuracy requirements. The framework is evaluated on an IoT-based plant disease classification task considering multiple deep learning architectures, federated aggregation strategies, and communication-round settings. Experimental results across 30 independent optimization runs show that the proposed approach explores only 11.11% of the search space, while consistently identifying solutions within 1% of the exhaustive-search optimum, with a mean optimality gap of only 0.056%. Index Terms—Hierarchical Federated Learning, Bayesian Optimization, Internet of Things, Smart Agriculture, Plant Disease Classification
I. I NTRODUCTION The agricultural sector is undergoing a major transformation, which is mainly driven by the combination of Internet of Things (IoT) technologies and deep learning (DL) models. This shift has already created new application scenarios such as data-driven crop monitoring, early disease detection and precision resource management [1], [2]. Within this context, the identification of plant diseases remains a critical challenge that undermines food security, since delayed interventions may result in substantial yield losses and economic damage [3]. A possible solution could reside within the recent advances in computer vision and convolutional neural networks, which
have achieved high accuracy in automated plant disease classification [4], [5]. However, the deployment of centralized DL pipelines in large-scale agricultural environments raises practical concerns. This mainly happens because continuous data transmission to remote cloud servers is impractical for resource-constrained IoT networks that operate under limited bandwidth, energy, and computational budgets [6]. Federated Learning (FL) has emerged as a compelling distributed paradigm that enables multiple IoT nodes to collaboratively train a shared model without exchanging raw data. Thus, this approach preserves data privacy and reduces communication overhead, compared with other distributed deployments [7], [8]. Several recent works have explored FL for agricultural applications, including crop disease classification [9], [10] and yield prediction [11]. These works demonstrate that FL can achieve competitive performance, while respecting the privacy and autonomy of individual farm sites. Despite these advances, the practical deployment of FL on heterogeneous IoT devices remains challenging, as system performance is highly sensitive to several configuration parameters. Such parameters include but are not limited to the backbone DL architecture, the model aggregation strategy (e.g., FedAvg [8], FedProx [12], FedAvgM [13]), the number of communication rounds and the number of participating devices [14]. In the majority of the existing literature, these design choices are selected through manual experimentation or via an exhaustive search. Unfortunately, both strategies are prohibitively expensive when each configuration evaluation requires end-to-end execution of the full FL pipeline. In this paper, we extend [15], [16] by introducing a constrained Bayesian Optimization framework for resource-aware Hierarchical Federated Learning (HFL) configuration in IoTbased plant disease classification. The proposed framework systematically explores deployment-specific configurations while accounting for resource and predictive-performance requirements. The main contributions of this work are summarized as follows:
We formulate HFL configuration as a constrained optimization problem that jointly considers architectural and training parameters under deployment-specific requirements. • We introduce a configurable weighted objective function that captures user-defined trade-offs among energy consumption, execution time, and predictive performance while enforcing explicit resource and performance constraints. • We integrate the spatial characteristics of the agricultural deployment into the framework to determine the required federation size. • We employ constrained Bayesian Optimization to jointly select the DL architecture, aggregation strategy, and number of communication rounds without exhaustively evaluating the configuration space. • We evaluate the proposed framework on a plant disease classification task in a resource-constrained IoT setting, demonstrating near-optimal configuration with a limited number of evaluations.
•
The remainder of this paper is organized as follows. Section II provides the background on Federated Learning and the considered deep learning architectures. Section III formulates the resource-aware HFL configuration problem, while Section IV presents the proposed constrained Bayesian Optimization approach. Section V presents the experimental evaluation. Finally, Section VI concludes the paper. II. BACKGROUND A. Federated Learning In Federated Learning (FL), every participating device uses its own private data to train a local model. A centralized controller collects and combines these individual contributions into a single global model and distributes it back to the devices [8]. In contrast to conventional single-server FL, HFL introduces intermediate aggregation layers between participating devices and the global server, enabling model aggregation closer to the data sources [17]. The aggregation strategy directly affects convergence and model quality. This work considers three strategies: (i) FedAvg [8], which computes a weighted average of local parameters; (ii) FedProx [12], which adds a proximal regularization term to limit client drift; and (iii) FedAvgM [13], which introduces server-side momentum to smooth successive global updates. B. Deep Neural Network Models In this work, we consider the following models: (i) EfficientNet-B0 [18]; (ii) ResNet-50 [19]; and (iii) MobileNetV3-Large [20]. These architectures exhibit different computational characteristics and provide a diverse set of backbone alternatives for evaluating the trade-offs between predictive performance and resource requirements. They therefore define the model-architecture dimension of the optimization search space.
III. R ESOURCE -AWARE HFL C ONFIGURATION P ROBLEM The deployment of HFL systems in resource-constrained IoT environments requires balancing predictive performance against energy consumption and execution time. To address this trade-off, we formulate the HFL configuration as a constrained optimization problem, where deployment requirements define the feasible configuration space. The framework incorporates user-defined preferences for energy consumption, execution time, and predictive performance, together with explicit resource and performance constraints. The federation size is determined by the spatial coverage requirements of the agricultural deployment, while the backbone architecture, aggregation strategy, and number of communication rounds constitute the optimization variables. These elements jointly define the resource-aware HFL configuration problem considered in the following sections. A. Search Space and Deployment Model A candidate HFL configuration is represented as x = (m, a, R),
X = M × A × {1, . . . , Rmax },
(1)
where m ∈ M denotes the selected backbone model, a ∈ A denotes the aggregation strategy, and R denotes the number of communication rounds. The number of participating devices is determined before optimization according to the farm geometry. Let Afarm denote the farm area and rd the effective sensing or communication radius of one device. Since circular coverage regions cannot cover an arbitrary agricultural area without overlap and boundary losses, a coverage-efficiency coefficient ρ ∈ (0, 1] is introduced. The required federation size is then estimated as Afarm , ρ = 0.8, (2) N= ρπrd2 where the adopted value assumes that 80% of the ideal circular region contributes to effective farm coverage. The value of ρ can be adjusted according to the geometry and coverage characteristics of a specific deployment [21]. Consequently, N is treated as a deployment parameter rather than a variable selected by the Bayesian optimizer. Let N0 and R0 denote the federation size and number of communication rounds used in the reference experiments. For each model–aggregator pair (m, a), the measured energy consumption E0 (m, a) and execution time T0 (m, a) are extended to different deployment conditions using firstorder scaling approximations. Motivated by resource-aware FL models, which characterize the overall computation and communication cost as accumulating across participating devices and communication rounds [14], [22], the energy consumption is approximated as N R E(x, N ) = E0 (m, a) , (3) N0 R0 and
T (x) = T0 (m, a)
R R0
where the user-defined coefficients λ1 , λ2 , and λ3 express the relative importance assigned to energy consumption, execution time, and predictive performance, respectively. To form a valid convex combination, they satisfy
.
(4)
The energy model accounts for changes in both federation size and number of communication rounds. The executiontime model depends only on R, under the assumption that local client operations are executed predominantly in parallel and that additional coordination overhead remains limited. To characterize the dependence of predictive performance on the number of communication rounds, a saturation function is fitted separately for each model–aggregator pair:
λi ≥ 0,
3 X
λi = 1.
(8)
i=1
The optimization is additionally restricted by the maximum allowable energy consumption and execution time, the minimum required predictive performance, and the admissible number of communication rounds: E(x, N ) ≤ Ebudget ,
F1 (x) = F1,0 (m, a)+[F1,∞ (m, a) − F1,0 (m, a)] 1 − e
−km,a R
(5) where F1,0 (m, a) denotes the initial performance, F1,∞ (m, a) the expected plateau, and km,a the corresponding convergence rate. These parameters are estimated from the round-level validation F1-score measurements recorded for each configuration. In the present experimental evaluation, however, direct measured F1 values are used whenever they are available. The fitted function is retained as a general round-dependent approximation for configurations or communication rounds for which direct measurements are not available. The upper search limit Rmax is determined from the experimentally available communication-round range. Although the fitted convergence curves can also be used to inspect the expected plateau behavior beyond the observed interval, round-level measurements in the present study are available for R = 1, . . . , 30. The experimental search space is therefore restricted to Rmax = 30, avoiding reliance on extrapolated performance values during the evaluation of the optimization method. B. Weighted Objective and Constraints Energy consumption and execution time are expressed in different physical units and may vary substantially across deployment conditions. They are therefore normalized directly by the budgets specified by the user, while the F1-score requires no additional scaling because it already lies within [0, 1]:
Ê(x, N ) =
E(x, N ) , Ebudget
T̂ (x) =
T (x) , Tbudget
F̂1 (x) = F1 (x).
(6) This formulation provides a direct interpretation of resource usage, since values of Ê or T̂ greater than one indicate that the corresponding deployment budget has been exceeded. The normalized quantities are combined into the scalar objective h i L(x) = λ1 Ê(x, N ) + λ2 T̂ (x) + λ3 1 − F̂1 (x) ,
(7)
,
T (x) ≤ Tbudget , F1 (x) ≥ F1,required ,
(9)
1 ≤ R ≤ Rmax . The budget-normalized terms determine the relative cost of feasible candidates, while the explicit constraints exclude configurations that violate the deployment requirements. C. Optimization Problem Let Ω ⊆ X denote the subset of configurations satisfying the constraints in (9). The configuration-selection problem is then expressed as x∗ = arg min L(x). x∈Ω
(10)
For a previously untested configuration, evaluating L(x) requires executing the corresponding HFL training process and obtaining its predictive and resource-related quantities. The configuration-selection problem can therefore be treated as an expensive black-box optimization task. In the present study, previously collected round-level measurements are used to emulate these expensive evaluations, allowing the proposed optimizer to be assessed under a controlled setting and directly compared with the optimum obtained from the complete discrete search space. The constrained Bayesian Optimization procedure used to perform this search is presented in the following section. IV. C ONSTRAINED BAYESIAN O PTIMIZATION Although the optimization problem in (10) is defined over a finite search space, exhaustively evaluating every candidate becomes increasingly expensive as additional architectures, aggregation strategies, communication-round settings, or deployment conditions are introduced. Each previously untested configuration may require the execution of the corresponding HFL training process together with the collection of predictive and resource-related measurements. Bayesian Optimization (BO) is therefore employed to guide the search toward promising feasible configurations while limiting the number of expensive evaluations [23]. The proposed procedure treats objective quality and constraint satisfaction separately. A Gaussian Process regression model approximates the scalar objective L(x), while a second
probabilistic model estimates the likelihood that a candidate satisfies the deployment constraints. The two models are combined through a constrained acquisition function that favors configurations expected to improve the current best feasible solution while maintaining a high probability of feasibility. A. Surrogate Models and Candidate Encoding
R−1 . Rmax − 1
(11)
(12)
where em and ea denote the one-hot vectors associated with the selected model and aggregation strategy, respectively. Given the set of already evaluated configurations Dt = {(zi , Li )}ti=1 , a Gaussian Process regression model is fitted to the observed objective values: L(z) ∼ GP (µ(z), k(z, z′ )) ,
(13)
providing, for each unevaluated candidate, a predictive mean µt (z) and standard deviation σt (z). In the implementation, a Matérn kernel is used for the objective surrogate. Constraint satisfaction is modeled independently. Each evaluated candidate is assigned the binary label ( 1, xi ∈ Ω, yi = (14) 0, xi ∈ / Ω, where Ω is the feasible set defined in (9). A Gaussian Process classifier is then trained on these labels to estimate pt (x) = P (x ∈ Ω | Dt ),
∆t (x) + σt (x)ϕ σt (x)
,
(15)
which represents the probability that a candidate satisfies the energy, execution-time, and predictive-performance requirements. B. Acquisition and Search Procedure The optimization begins with a small set of unique randomly selected configurations. After these initial evaluations, the objective surrogate and feasibility model are updated using all observations collected so far. If no feasible configuration has yet been observed, candidate selection is driven by the estimated probability of feasibility until the first feasible solution is identified. Candidate selection is based on Expected Improvement (EI), which quantifies the expected reduction relative to the best feasible objective value observed at iteration t. For minimization, EI is computed as
(16)
where (17)
Φ(·) and ϕ(·) denote the standard normal cumulative and probability density functions, respectively, and ξ controls the exploration–exploitation trade-off. To account explicitly for the deployment constraints, EI is weighted by the estimated probability of feasibility, following the constrained BO formulation in [24]: αt (x) = EIt (x) pt (x).
The resulting numerical representation is z(x) = [em , ea , R̃],
∆t (x) EIt (x) = ∆t (x)Φ σt (x)
∆t (x) = Lbest − µt (x) − ξ,
The search variables contain both categorical and numerical components. For a candidate x = (m, a, R), the backbone model m and aggregation strategy a are represented through one-hot encoding, while the communication round is normalized to the interval [0, 1] as R̃ =
(18)
At each iteration, the acquisition function is evaluated over the set of configurations that have not yet been tested, and the next candidate is selected according to xt+1 = arg max αt (x). x∈X \Dt
(19)
Only the selected candidate is then evaluated using the actual HFL objective and constraints, after which both probabilistic models are updated and the process is repeated until the predefined evaluation budget is reached. Restricting acquisition evaluation to previously unseen candidates also guarantees that no configuration is evaluated more than once. In the present study, six unique configurations are used for initialization and the overall optimization budget is limited to 30 evaluations. V. E XPERIMENTAL E VALUATION A. Optimization Setup and Ground-Truth Benchmark The proposed optimization framework was evaluated using the nine HFL model–aggregator combinations considered in the previous experiments, with the number of communication rounds restricted to R ∈ {1, . . . , 30}. The resulting discrete search space contains |X | = 3 × 3 × 30 = 270
(20)
candidate configurations. The reference measurements correspond to N0 = 10 participating clients and R0 = 30 communication rounds. For the deployment scenario considered in the optimization experiments, we assume a farm area of Afarm = 10 000 m2 and an effective device coverage radius of rd = 20 m. With ρ = 0.8, the spatial coverage model in (2) results in N = 10, matching the federation size used in the reference experiments. Consequently, no additional scaling with respect to the number of participating clients is introduced in the present evaluation, while the proposed formulation remains applicable to deployments with different federation sizes. The optimization parameters were set to
Ebudget = 20 Wh,
Tbudget = 400 s,
F1,required = 0.80, (21)
with objective weights (λ1 , λ2 , λ3 ) = (0.4, 0.2, 0.4).
(22)
Energy consumption and execution time at intermediate communication rounds were obtained through (3) and (4), while the directly measured validation F1-score was used for each available round. Since measurements were available for the complete search space, an exhaustive evaluation of all 270 candidates was performed once as an offline reference. This exhaustive search is not part of the proposed optimization procedure, but is used exclusively to establish the true optimum and quantify the performance of Bayesian Optimization. Of the 270 candidate configurations, 141 satisfied all deployment constraints. The globally optimal feasible solution was EfficientNet-B0 with FedAvg at R = 4, yielding E = 1.556 Wh,
T = 40.156 s,
F1 = 0.800326, (23)
with an objective value of L∗ = 0.131068.
(24)
The second-best configuration was ResNet-50 with FedProx at R = 2, with L = 0.131266, corresponding to a relative difference of only 0.152% from the global optimum. B. Bayesian Optimization Results To evaluate robustness with respect to initialization, the constrained Bayesian Optimization procedure was repeated for 30 independent random seeds. Each run was initialized with six unique randomly selected configurations and was limited to a total budget of 30 evaluations. Therefore, each optimization run evaluated only 30 × 100 = 11.11% 270 of the complete configuration space.
(25)
TABLE I P ERFORMANCE OF CONSTRAINED BAYESIAN O PTIMIZATION OVER 30 INDEPENDENT RUNS . Metric Search-space size Evaluations per run Feasible-run rate Exact global optimum Within 1% of optimum Mean optimality gap Median optimality gap Maximum optimality gap Median best iteration
while the exact global optimum was recovered in 63.33% of the cases. More importantly, every run terminated with a solution within 1% of the exhaustive-search optimum. The mean relative optimality gap was only 0.056%, the median gap was 0%, and the maximum observed gap was 0.152%. In all runs in which the exact optimum was not selected, the final solution corresponded to the second-best configuration identified by exhaustive evaluation. The convergence behavior is shown in Fig. 1. The median best-so-far objective decreases rapidly as additional configurations are evaluated and approaches the exhaustive-search optimum after approximately 15–20 evaluations. After only 15 evaluations, corresponding to 5.56% of the complete search space, the median optimality gap was already 0.152%, while 56.67% of the runs had reached a solution within 1% of the optimum. This proportion increased to 80.00% after 20 evaluations, 86.67% after 25 evaluations, and 100% at the final budget of 30 evaluations. The final best solution of each run was first identified at a median iteration of 15.5, corresponding to approximately 5.74% of the complete search space. These results indicate that the proposed constrained Bayesian Optimization procedure can reliably concentrate the search around the globally optimal region while requiring only a small fraction of the evaluations needed by exhaustive search.
Result 270 30 (11.11%) 100% 63.33% 100% 0.056% 0.000% 0.152% 15.5
Table I summarizes the final optimization performance over all 30 runs. A feasible solution was identified in every run,
VI. C ONCLUSIONS This paper presented a constrained Bayesian Optimization framework for efficient HFL configuration in resourceconstrained IoT environments for plant disease classification. The framework jointly optimizes the learning architecture, aggregation strategy, and communication rounds under energy, execution-time, and predictive-performance constraints. Experimental results show that only 11.11% of the search space is explored, while all runs identified solutions within 1% of the exhaustive-search optimum. These results demonstrate the potential of constrained Bayesian Optimization for efficient and resource-aware HFL deployment in smart agricultural IoT environments. ACKNOWLEDGMENT This work was funded under the COIN-3D project, which has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No. 101159667. R EFERENCES [1] S. Wolfert, L. Ge, C. Verdouw, and M.-J. Bogaardt, “Big data in smart farming – a review,” Agricultural Systems, vol. 153, pp. 69–80, 2017. [2] O. Elijah, T. A. Rahman, I. Orikumhi, C. Y. Leow, and M. N. Hindia, “An overview of internet of things (iot) and data analytics in agriculture: Benefits and challenges,” IEEE Internet of Things Journal, vol. 5, no. 5, pp. 3758–3773, 2018. [3] V. K. Vishnoi, K. Kumar, and B. Kumar, “Plant disease detection using computational intelligence and image processing,” Journal of Plant Diseases and Protection, vol. 128, no. 1, pp. 19–53, 2021.
Fig. 1. Convergence of constrained Bayesian Optimization over 30 independent runs. The solid curve represents the median best feasible objective observed up to each evaluation, the shaded region denotes the interquartile range, and the dashed line indicates the global optimum obtained from exhaustive evaluation of the complete search space.
[4] A. Upadhyay, N. S. Chandel, K. P. Singh, S. K. Chakraborty, B. M. Nandede, M. Kumar, A. Subeesh, K. Upendar, A. Salem, and A. Elbeltagi, “Deep learning and computer vision in plant disease detection: a comprehensive review of techniques, models, and trends in precision agriculture,” Artificial Intelligence Review, vol. 58, no. 3, p. 92, 2025. [5] G. Delnevo, R. Girau, C. Ceccarini, and C. Prandi, “A deep learning and social iot approach for plants disease prediction toward a sustainable agriculture,” IEEE Internet of Things Journal, vol. 9, no. 10, pp. 7243– 7250, 2022. [6] P. K. Kashyap, S. Kumar, A. Jaiswal, M. Prasad, and A. H. Gandomi, “Towards precision agriculture: Iot-enabled intelligent irrigation systems using deep learning neural network,” IEEE Sensors Journal, vol. 21, no. 16, pp. 17 479–17 491, 2021. [7] E. T. Martı́nez Beltrán, M. Q. Pérez, P. M. S. Sánchez, S. L. Bernal, G. Bovet, M. G. Pérez, G. M. Pérez, and A. H. Celdrán, “Decentralized federated learning: Fundamentals, state of the art, frameworks, trends, and challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 4, pp. 2983–3013, 2023. [8] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017, 20-22 April 2017, Fort Lauderdale, FL, USA, ser. Proceedings of Machine Learning Research, A. Singh and X. J. Zhu, Eds., vol. 54. PMLR, 2017, pp. 1273–1282. [Online]. Available: http://proceedings.mlr.press/v54/mcmahan17a.html [9] P. Hari and M. P. Singh, “Adaptive knowledge transfer using federated deep learning for plant disease detection,” Comput. Electron. Agric., vol. 229, no. C, Feb. 2025. [10] S. Behera, N. Padhy, R. Panigrahi, and S. Kumar Kuanar, “Crop disease prediction using deep learning in a federated learning environment: Ensuring data privacy and agricultural sustainability,” Procedia Computer Science, vol. 254, pp. 137–146, 2025, international Conference on Digital Sovereignty (ICDS). [11] S. Bera, T. Dey, A. Mukherjee, and D. De, “Flag: Federated learning for sustainable irrigation in agriculture 5.0,” IEEE Transactions on Consumer Electronics, vol. 70, no. 1, pp. 2303–2310, 2024. [12] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” 2020. [Online]. Available: https://arxiv.org/abs/1812.06127 [13] T.-M. H. Hsu, H. Qi, and M. Brown, “Measuring the effects of nonidentical data distribution for federated visual classification,” 2019. [14] S. Wang, T. Tuor, T. Salonidis, K. K. Leung, C. Makaya, T. He, and
K. Chan, “Adaptive federated learning in resource constrained edge computing systems,” IEEE Journal on Selected Areas in Communications, vol. 37, no. 6, pp. 1205–1221, 2019. [15] A. Papanikolaou, A. Tziouvaras, G. Floros, A. Xenakis, and F. Bonsignorio, “Distributed deep learning in iot sensor network for the diagnosis of plant diseases,” Sensors, vol. 25, no. 24, 2025. [16] A. Papanikolaou, A. Tziouvaras, P. Stoikos, A. Xenakis, S. A. Puthiya Parambath, G. Floros, E. Zereik, I. Petrovic, and F. Bonsignorio, “Performance and energy trade-off analysis of hierarchical federated learning for plant disease classification,” in 2026 IEEE Engineering Reliable Autonomous Systems (ERAS), 2026, pp. 96–101. [17] L. Liu, J. Zhang, S. Song, and K. B. Letaief, “Client-edge-cloud hierarchical federated learning,” in ICC 2020 - 2020 IEEE International Conference on Communications (ICC), 2020, pp. 1–6. [18] M. Tan and Q. V. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” 2020. [Online]. Available: https://arxiv.org/abs/1905.11946 [19] K. He, X. Zhang, S. Ren, and J. Sun, “ Deep Residual Learning for Image Recognition ,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2016, pp. 770–778. [20] A. Howard, M. Sandler, B. Chen, W. Wang, L.-C. Chen, M. Tan, G. Chu, V. Vasudevan, Y. Zhu, R. Pang, H. Adam, and Q. Le, “ Searching for MobileNetV3 ,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Los Alamitos, CA, USA: IEEE Computer Society, Nov. 2019, pp. 1314–1324. [21] H. M. Ammari, “A computational geometry-based approach for planar k-coverage in wireless sensor networks,” ACM Trans. Sen. Netw., vol. 19, no. 2, Feb. 2023. [22] Z. Yang, M. Chen, W. Saad, C. S. Hong, and M. Shikh-Bahaei, “Energy efficient federated learning over wireless communication networks,” IEEE Transactions on Wireless Communications, vol. 20, no. 3, pp. 1935–1949, 2021. [23] J. Snoek, H. Larochelle, and R. Adams, “Practical bayesian optimization of machine learning algorithms,” Advances in neural information processing systems, vol. 25, 2012. [24] J. R. Gardner, M. J. Kusner, Z. E. Xu, K. Q. Weinberger, and J. P. Cunningham, “Bayesian optimization with inequality constraints,” in Proceedings of the 31st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 32, no. 2, 22–24 Jun 2014, pp. 937–945.