arXiv:2609.10151v1 [cs.DC] 9 Sep 2026
CEDD-optimizer: Enabling Cost-Efficient Dataset Distillation on Geographically Distributed Edge Systems Dai Liu
Eishi Arima
Martin Schulz
Technical University of Munich Garching, Germany [email protected]
Technical University of Munich Garching, Germany [email protected]
Technical University of Munich Garching, Germany [email protected]
Abstract—Centralized learning is a fundamental paradigm in modern AI, where data are typically collected from distributed edge devices and subsequently aggregated at a central host for model training. However, the overall training pipeline is often bottlenecked by the substantial communication overhead incurred during the data collection. Dataset Distillation (DD), benefiting from its remarkable compression ratio, has emerged as a leading dataset compression technique, making it particularly attractive for centralized learning on distributed data. However, while existing studies have demonstrated DD’s overwhelming compression ratio, its cost efficiency in non-uniform edge environments has been largely overlooked. As edge devices are often distributed geographically in the real world, their energy and data transfer prices are often non-uniform. At the same time, several key hyperparameters in DD (e.g., the target compression ratio and the number of distillation steps) affect the energy and data transfer overhead considerably as well as the training quality (or the test accuracy of downstream training tasks), requiring careful tuning both locally and globally. In this work, we propose Cost-Efficient Dataset Distillation optimizer (CEDDoptimizer), a hyperparameter tuning framework for cost-efficient distributed DD. Our framework aims at minimizing the total cost under a constraint for training quality by optimizing the hyperparameter settings across edge devices while being aware of the environmental non-uniformity. Our framework relies on two key modules: CEDD-calibrator and CEDD-solver. The CEDDcalibrator identifies parameters in our energy and training quality modeling — the former is detected by an offline calibration, while the latter is estimated online during our three-step tuning scheme. Building upon these models, the CEDD-solver deals with the cost minimization problem to steer and improve the distributed DD workflow. Our thorough experiments across various image datasets show that our approach achieves up to a 20.8x improvement over the baseline DD method under the same quality constraint. Index Terms—Dataset Distillation, Heterogeneous Edge Environment, Data Compression
I. I NTRODUCTION Deep learning [1] has been widely adopted across a broad range of applications, all of which rely on training deep neural networks with large amounts of data. A common training paradigm is centralized learning [2], where a training dataset, typically collected from edge devices distributed across different locations, is aggregated and used to train a model at a central host. Real-world applications rely on this paradigm,
e.g., for edge-cloud defect detection in manufacturing [3], privacy-preserving federated histopathology across hospitals [4], distributed crop monitoring in agriculture [5], [6], managing heterogeneous edge devices via centralized slimmable networks [7]. However, the training pipelines are often costly and constrained by communication overhead when transferring large datasets from edge devices to the host. Traditional lossy and lossless compression techniques [8]– [10] do not effectively resolve the communication burden. On one hand, lossless compression techniques alone do not mitigate the data transfer bottleneck for deep learning tasks due to their limited compression ratios. On the other hand, the deep neural network training process relies on subtle and fine-grained details in the data for effective learning, whereas lossy compression methods, which achieve higher compression ratios, inevitably sacrifice these details, resulting in poor training quality. Dataset Distillation (DD) [11]–[13] has emerged as a promising alternative to traditional dataset compression techniques. Unlike traditional training tasks, DD utilizes network gradients to train a smaller synthetic dataset. Compared to traditional compression methods, DD offers the following major advantages: (i) very high compression ratio [11]–[17]; (ii) marginal accuracy loss [15], [16]; (iii) data anonymization properties by synthesizing data from the original dataset [18]– [20]; and (iv) no decompression required, thus a significant overhead reduction for downstream training tasks including continual learning [21]–[26]. These properties collectively address the aforementioned obstacles, making DD an excellent fit for efficient centralized learning scenarios. Despite its substantial benefits, it relies on a costly hyperparameter tuning step, which is a major challenge when deploying DD in distributed edge environments. This, however, has been largely overlooked in the literature. In particular, two key hyperparameters significantly affect both test accuracy and operational cost: the target compression ratio and the distillation loop iteration count. However, in existing DD methods [13], [27], [28], these parameters are manually tuned without accounting for real-world environments, where systems are typically non-uniform, particularly in terms of data transfer and electricity costs, depending on the geographical locations of
Conventional Solution:
Our Solution:
CEDD-optimizer
! Price Variation-aware Setup
$: Electricity & Data Transfer Price $$$
! Costly Region
Edge1
DD IC (β) = 30
Compressed to DPC (α) = 25
DD IC (β) = 11
Edge1
Edge2
$
DD IC (β) = 30 Edge3
Compressed to DPC (α) = 25
Neural Network Training Host
$$
DD IC (β) = 30
Edge2
$$$
αs & βs Compressed to DPC (α) = 15
$
DD IC (β) = 32
Compressed to DPC (α) = 27
Neural Network Training Host
Edge3 $$
Compressed to DPC (α) = 25
!Uniform Setup ! Unware of Geographical Price Variation
DD IC (β) = 25
! Total Cost Reduction
Compressed to DPC (α) = 20
! The Same Training Quality
Fig. 1. The Basic Concept of Our Work: Price Variation-aware Tuning for Data Per Class (DPC, α) and Iteration Count (IC, β)
edges. Therefore, merely deploying conventional DD methods results in a suboptimal hyperparameter setup (e.g., assigning a high DD compute load to an energy-intensive region) when facing a key optimization challenge: minimizing the total operational cost while satisfying a requirement for model quality. In this work, we focus on this significant optimization challenge and present Cost-Efficient Dataset Distillation optimizer (CEDD-optimizer), a novel hyperparameter tuning framework, with a particular focus on distributed DD with non-uniform device setups. The basic concept is illustrated in Fig. 1. Our work targets a centralized learning pipeline in which each edge device applies DD to its local data to produce a compact, distilled dataset. These distilled datasets are then aggregated at a central training host to train a deep neural network. Our CEDD-optimizer optimizes the two key DD hyperparameters for each edge node separately: (i) distilled data per class (DPC, α), which determines the volume of transferred data, and (ii) iteration count (IC, β), which governs local computation and energy consumption. To this end, it coordinates two key components: an offline CEDD-calibrator and an online CEDD-solver. The former efficiently identifies the unknowns of two predictive models used to estimate energy consumption (offline) and test accuracy (online). The CEDD-solver then employs these predictive models to solve the optimization problem that minimizes total system cost (including communication, energy, and storage) for a given target test accuracy for downstream deep learning tasks. By selecting the optimal pair (α, β) for each edge device separately, our CEDD-optimizer provides tailored configurations that align the quality of the aggregated distilled dataset with the system cost-efficiency across non-uniform, geographically distributed deployments. The following are the major contributions of this paper: • Identification of the new challenge: To the best of our knowledge, this is the first work to identify the challenge of hyperparameter tuning in DD under non-uniform, geographically distributed edge environments, accounting for the total economic cost arising from heterogeneous networks and energy prices across geographical locations. • Formulation and modeling: We formulate the hyperparameter tuning challenge as a formal optimization problem in a concrete mathematical format. We construct simple predictive models to estimate the test accuracy of deep learning tasks
trained with the distilled dataset, the energy consumption of DD on the edges, and the volume of the distilled dataset as functions of DD’s hyperparameters. • The CEDD-optimizer framework: We offer a cost-aware hyperparameter tuning framework for distributed DD that coordinates the CEDD-calibrator and the CEDD-solver. The former efficiently identifies the model coefficients offline or online, while the latter solves the optimization problem using the models. • Thorough evaluation: We thoroughly evaluate and validate the proposed models and optimizer in diverse configurations, using four different edge devices, five benchmark datasets, and a wide range of geographically distributed price data. Our results demonstrate that CEDD-optimizer consistently achieves significant cost reductions while maintaining competitive accuracy, and we further analyze its scalability and trade-offs under varying system conditions. Find our code in https://github.com/NiaLiu/CEDD-optimizer.git. II. BACKGROUND AND R ELATED W ORK A. Dataset Distillation Dataset Distillation (DD) was first proposed by Wang et al. in 2018 [11] as an efficient data compression method specialized for model training in deep learning tasks. Various followup studies [12], [14], [15], [29], [30] have proven that DD significantly outperforms conventional lossy or lossless dataset compression methods, including the Coreset Selection [12], [14], [30] and others [31]–[35]. Recent DD algorithms can be split into iterative methods and non-iterative methods. The iterative methods are considered variants of DC (Dataset Condensation) [12]. It typically relies on iterative bi-level optimization, such as aligning feature distributions [29], addressing classes’ miss-alignment [29], [36], capturing classes’ differences [37], aligning layer-wise features [14], embedding differentiable siamese augmentation [38], applying a kernel-based metalearning framework [39], and adopting soft labels [40]–[42], etc. [43]–[46]. Further, non-iterative methods are often one shot [47]–[51]. In this paper, we focus on iterative methods, which are generally costly but provide good results. For generality, we choose the original DC [12] as our baseline. Our approach is not specific to DC and is generally applicable to other iterative DD methods, as we target key parameters commonly available in any DD method. Overall, our approach is the first to introduce a geographical location-based hyperparameter tuning methodology in DD to minimize the total cost with marginal accuracy loss for edge deployments. B. Edge Computing The concept of edge computing has gained momentum ever since it was proposed in the 2000s [52], and various studies have focused on big data-driven AI applications in edge computing. Several recent studies target cost management for edge computing or the edge-to-cloud continuum [53]– [56]. Lu et al. [57] propose a hybrid method that combines lossy and lossless compression for edge computing, and
Gateway
.
. ..
Cloud Storage Data Cost Data Cost
..
. ..
M Data in Total
.....
Training
Original Dataset Neural Network
Dataset Distillation
Benefit2: Data Benefit1: Outstanding Benefit3: Efficient Anonymization Compression Ratio Training Data per α×K (<< M) Data in Total Class: α ..... Training Classes (Total #: K) Distilled Synthetic Dataset
Similar Accuracy
DD Energy Cost
Network Data Cost DD
Hyper Params
Energy Cost
Training Host (Centralized Learning) Data Cost Hyper Params Energy Cost
DD
Hyper Params
Edge Server
Neural Network
Fig. 2. Concept and Benefits of DD: Replacing the Original Dataset with a Synthetic One for Efficient Training Conpression Ratio
Energy Consumption
*
*
*
Fig. 3. Comparisons of Compression Ratio and Energy among Various Compression Methods
Wu et al. [58] propose a CNN-based encoder and decoder. Our work introduces DD as a promising alternative. There are several tools for service deployment in edge computing. For example, Rosendo et al. [59] provide E2Clab. They later extend the framework to support an optimization functionality [60] and enable efficient provenance capture in IoT/Edge [61] environments. Some studies spot on environmental constraints in the edge-to-cloud continuum, including (1) network bandwidth constraint and network slicing optimization [62] and (2) energy constraint for battery-powered or energy-harvesting devices [63]–[65]. Others target heterogeneous edge computing environments [66], [67]. It is worth noting that our work is orthogonal to decentralized learning methods, such as Federated Learning (FL) [68], [69]. While DD can be integrated into FL frameworks to enhance communication efficiency and data privacy [13], [27], [28], [70]–[72], this integration lies beyond the scope of this study. In this paper, we focus instead on improving the efficiency of centralized learning pipelines, where DD serves as the core mechanism for data compression and quality/cost optimization. III. I NTRODUCING DATASET D ISTILLATION TO E DGE S YSTEMS A. Motivation: High Compression Ratio of DD Fig. 2 depicts the overall concept of DD. DD converts an original dataset (M data in total) into a much smaller number of synthetic distilled data. Subsequently, the distilled dataset is used to train a given neural network. As the synthetic distilled
Fig. 4. Our Target System Architecture and DD Scenario with Data / Energy Cost Management via Hyperparameter Tuning
dataset conceals the training information of the original dataset, the neural network trained on the distilled dataset can achieve inference accuracy similar to that of the original dataset. DD has been applied to classification tasks in the literature, and the number of data per class (or DPC) is a key hyperparameter that determines the overall data compression ratio. Let M , K, and α be the total size of the original dataset, the number of classes in the dataset, and the DPC in the distilled dataset, respectively. Then, the compression ratio is simply denoted as M/(α · K). Therefore, the compression ratio in DD is fully controllable by the hyperparameter α. Fig. 3 compares the compression ratio and energy consumption among DD and other existing lossless or lossy compression mechanisms: PNG [73], WebP, int8Q (an int8 quantization) [74], CAE (ConvAE, a classic convolutional autoencoder) [75], and DAE (DenseAE, a fully connected autoencoder) [76]. As shown in the left graph of Fig. 3, DD achieves an outstanding compression ratio, outperforming the others by orders of magnitude improvement. In this evaluation, we used a simple three-layer convolutional neural network as the baseline model for both DD and CAE. We set the test accuracy thresholds for training tasks to 97% and 60% for the MNIST [77] and CIFAR10 [78] datasets, respectively. As for the energy consumption, although DD requires a noticeable energy overhead for its compression procedure, which is also controllable by tuning hyperparameters, it can significantly reduce the overhead for downstream model training using the distilled dataset. As a consequence, the total energy cost of DD is the smallest in this experiment. In this energy evaluation, we use the same platform (NVIDIA Jetson Xavier NX) for all computations, which will be described later in Section VI. B. Application Scenario Fig. 4 illustrates the scenario we assume in this paper. We target the following computing continuum: (1) data are generated on end devices geographically distributed across the world; (2) each edge server receives the data from local end devices and applies DD to the collected dataset; (3) the distilled dataset is aggregated and stored on the cloud storage; (4) the training host utilizes the distilled dataset to train a neural network model; and (5) the trained model is stored on the cloud storage. Note, this scenario is based on centralized
0.5
Electricity Cost [$/kWh]
0.45 0.4 0.35 0.3 0.25 0.2 0.15 0.1 0.05 0
Bermuda CaymanIslands Denmark Germany Czech… France Jamaica Lithuania Latvia Uruguay Portugal Spain Mali NewZealand Aruba Gabon Rwanda Senegal SriLanka Romania Chile Norway Macau Namibia Moldova Serbia SouthKorea Madagascar Eswatini Armenia Lesotho Ecuador Indonesia Cameroon China Belarus Suriname Georgia Afghanistan SaudiArabia Kazakhstan Azerbaijan Algeria Laos Zambia Iraq Angola Sudan Ethiopia
learning, commonly used in a wide variety of services [79]. We also account for storage costs in our experiments to cover the potential of continual learning [21]–[26]. In the target scenario, we use CEDD-optimizer to tune the hyperparameters of DD distributed across edges to manage the economic cost encompassing data transfer/storage and electricity for DD, which are highly dependent on the geographical locations of edge devices.
Country
C. Leveraging Hyperparameters in DD
α: # of Data per Class (DPC) Small Dataset
Low
α
High
High Quality
β: Iteration Count (IC) High
α
Low
High
β
Low
Low DD Energy
Fig. 5. Trade-off Relationships in DD among Dataset Size, Quality, and Energy Overhead
On one hand, DD is typically implemented as an iterative solver, and thus the execution time and energy consumption are linear functions of the IC (or β) that determines the count of outermost loop. On the other hand, the computational and energy overheads of DD per loop are determined by the number of data per class (DPC or α). In particular, they can be modeled by a quadratic function of DPC α. This is because the DD main loop processes α images x α times per iteration, resulting in a computational complexity of O(α2 ) per iteration. We model and verify this later in Section IV-B1 and VI-B1. D. Impact of Geographical Locations on Pricing We target the DD scenario illustrated in Fig. 4 where end devices and edge devices are geographically distributed at a global scale while being aware of both the total operating cost and the test accuracy of deep learning tasks. This global deployment results in variations of electricity cost for edge devices and data transfer pricing per size. Several recent studies also address cost management in edge computing or the edge-to-cloud continuum; however, they do not consider DD or the price variations of edge energy and data transfer simultaneously [53]–[56]. For example, the study by Huang et al. [54]
Fig. 6. Electricity Cost by Country in 2024 including 144 Countries (Remade from [80]) Cost per Data Transfer Size [$/GB] Far_East2
0.06
0.09
Far_East1
0.09
0.041
0.09
0.15 0.1132 0.1132 0.17 0.17columnheaders 0.042 0.042 ’sitelink-price’ matrix rowheaders 0.15 0.1107 0.1107 0.16 0.17 0.041 0.041
South_America 0.1107
0.15
0.16
0.18
0.19
0.16
0.11
0.17
0.17
0.17
Southern_Africa
0.11
0.17
0.14
0.19
0.11
0.16
0.16
0.16
0.17
0.11
0.19
0.14
0.042
0.19
0.19
0.1107 0.1107 0.1132
0.103
0.045
0.14
0.14
0.18
0.1107 0.1107 0.1132
North_America2
0.03
0.02
0.11
0.085
0.13
0.11
0.15
0.09
0.09
0.09
Middle_East
0.1
0.11
0.09
0.103
0.19
0.17
0.16
0.15
0.15
0.15
North_America1 0.0282
0.02
0.11
0.085
0.13
0.11
0.15
0.09
0.09
0.09
Europe
0.0282
0.1
0.085
0.13
0.11
0.1107
0.09
0.09
0.06
a
as t1 _E
_E
Fa r
Fa r
ic
0.1
0.05
0
So
ut
h_
A
m
_A rn he
ut So
er
fri
an
ia
sia A
ce O
h_ ut
m _A
N
or
th
M
So
er
_E
ic
as
t
a1 ic
id d
le
er
ro
m
Eu
_A th or N
0.15
as t2
0.13 0.085
ca
0.13 0.085
a2
Oceania
0.02
0.2
0.042
South_Asia
pe
Source
We leverage the key hyperparameters that govern the quality of the distilled dataset and the DD overhead. The objective is to minimize the cost of data transfer/storage and energy consumption at the edges while meeting a target quality requirement. As a quality metric, we use the test accuracy of downstream deep learning tasks trained with the distilled dataset, as it reflects the quality of the dataset and matters for the users as well. To this end, we deal track and adjust the two following parameters across all edge systems: (1) the number of data per class (or DPC) denoted as α and (2) the upper bound of iteration count for the outermost loop of DD (or IC) represented as β. Fig. 5 illustrates the trade-off relationships governed by these two types of hyperparameters. Both parameters affect the test accuracy of downstream deep learning tasks and the energy (or computational) overhead to perform DD on the edge, whereas the dataset size is affected only by the DPC (see also Section III-A).
Destination
Fig. 7. Data Transfer Cost (or SiteLink Rate) by Location in AWS Direct Connect (Remade from [81])
considers the costs of energy and communication, but does not account for their variations. To the best of our knowledge, our work is the first to target the combination of the geographical price variations and the application of DD to edge devices with hyperparameter tuning. Fig. 6 lists the electricity price [$/kWh] by country in 2024, encompassing 144 countries [80], while Fig. 7 demonstrates the communication cost per combination of source and destination locations [$/GB] in a global private network infrastructure service, in this case AWS direct connect [81], [82] as an example. In Fig. 6, the X-axis lists countries sorted by cost in descending order, whereas the Y-axis represents the electricity cost. As shown, the variation is significant by orders of magnitude. As for communication costs, we show the operating cost per data transfer when using a global private network provided by AWS [82] as an example. It enables both edge-toedge and edge-to-cloud communications in a direct and private fashion, while charging in a pay-as-you-go policy [81], [82]. As shown in Fig. 7, the communication cost varies by an order of magnitude depending on the geographical locations of edge devices. IV. F ORMULATION AND M ODELING A. Problem Formulation Fig. 8 illustrates the overall problem we address with this study. We target the system that comprises N edge nodes and one training host that performs the downstream learning task using the collected dataset from the edge nodes. On the ith edge
Control Params: α = [α1, ..., αN], β = [β1, ..., βN] Minimize Total Cost: min Cost(O, α, β, Ce, Cd, Cs) = min ΣiCieEi+ΣiCidSi+CsΣiSi DD Hyper α1 β1 Params:
Training Host
B. Cost and Accuracy Modeling
Meet Target Quality: Acc(D=DD(O, x)) ≥ Atrg C1d·S1 . . . . . Cid·Si α i βi
Data Transfer Costs
.....
CNd·SN
αN βN
DD DD DD D1 ON DN Oi Di α1,β1 αN,βN αi,βi ..... . . . . . S1=Size(D1) SN=Size(DN) Si=Size(Di) Orig. Dataset + EN=Edd(ON, αN, βN) Ei=Edd(Oi, αi, βi) Cost Params: E1=Edd(O1, α1, β1) O1, C1e, C1d ON, CNe, CNd Edge1 Edgei EdgeN Oi, Cie, Cid O1
→
→
→
In order to optimize the hyperparameter setup (α, β) by solving the problem formulated in Section IV-A, both the cost and accuracy functions (Cost() and Acc()) need to be known. In this section, we present our predictive modeling. 1) Cost Modeling: As presented in Fig. 8, the total cost is broken down into (1) energy cost, (2) data transfer cost, and (3) storage cost in this study. We model the cost function by using three dot-product terms as follows:
Fig. 8. Problem Overview: Minimizing the Total Cost (Cost) for a Target Test Accuracy (Acc) by Leveraging Hyperparameters across Edges (α = [α1 , · · · , αN ], β = [β1 , · · · , βN ])
Cost(O, α, β, Ce , Cd , Cs ) = =
Ce · E + Cd · S + Cs 1 · S X (Cie Ei + Cid Si + Cs Si )
(5)
1≤i≤N
where E = [E1 , · · · , EN ], S = [S1 , · · · , SN ], node, the hyperparameters αi (or DPC) and βi (or IC) need to be set accordingly, depending on given inputs or requirements. 1 = [1, · · · , 1] (6) The following parameters are associated with the ith node: In Eq. (5), these three dot products Ce · E, Cd · S, and Oi , Cie , and Cid . Here, Oi is the original dataset collected at the node, whereas Cie or Cid are the cost parameters that Cs 1 · S represent the energy cost to apply DD, the data transfer represent the price per energy [$/J] or the price per transferred price of distilled dataset, and the storage consumption cost, data volume [$/B], respectively. As mentioned in Section III-D, respectively. Eq. (6) lists the DD energy (E) and the data these cost parameters can vary significantly by location. Once volume to transfer (S) for N edge devices. Here, Ei denotes the hyperparameters αi and βi are set, the DD is applied to the DD energy consumed at the ith edge, while Si is equal to Oi to generate a distilled dataset Di . The aggregated distilled the size of the distilled dataset Di generated at the ith edge datasets D = [D1 , D2 , · · · , DN ] are used in the downstream node, both functions of Oi , αi , and βi . Next, we model the first term, i.e., the energy overhead of training task. Thus, its test accuracy Acc() is a function of the DD on edge devices empirically: aggregated distilled datasets D. In this study, we deal with the following optimization: minimizing the total cost Cost() for a Ei = Edd (Oi , αi , βi ) given targeted accuracy Acc() ≥ Atrg . To this end, the tuning = (d1,i αi2 + d2,i αi + d3,i ) · βi + d4,i (7) procedure determines α = [α1 , · · · , αN ], β = [β1 , · · · , βN ], tailored for the set of original datasets O = [O1 , O2 , · · · , ON ] As mentioned in Section III-C, DD is typically implemented and given price parameters including (1) the edge energy costs as an iterative solver. Consequently, the execution time and e per Joule Ce = [C1e , C2e , · · · , CN ], (2) the data transfer costs per Byte between the training host and edge devices Cd = d TABLE I [C1d , C2d , · · · , CN ], and (3) the cloud storage cost per Byte Cs . N OTATIONS OF VARIABLES /F UNCTIONS The optimization problem is formulated as follows: In :
O, Ce , Cd , Cs
Out : α = [α1 , · · · , αN ], β = [β1 , · · · , βN ]
Variable
Remarks
N, M, K, V
# of edge devices, the total # of original data across edges, # of classes in the classification task, the size per data DPC (Data Per Class) on the ith edge (αmin ≤ αi ≤ αmax ) IC (Iteration Count) on the ith edge (βmin ≤ βi ≤ βmax ) The vector of αi , βi : α = [α1 , · · · , αN ], β = [β1 , · · · , βN ]
min
Cost(O, α, β, Ce , Cd , Cs )
(1)
αi βi α, β
s.t.
Acc(D = DD(O, α, β)) ≥ Atrg
(2)
O
αmin ≤ αi ≤ αmax = M/(N · K · Rtrg ) (3)
D
βmin ≤ βi ≤ βmax (1 ≤ ∀ i ≤ N )
(4)
The objective is to minimize Cost() that encompasses the costs of data transfers, edge energy, and cloud storage, which is a function of O, α, β, Ce , Cd , and Cs . Eq. (3) and Eq. (4) show the ranges of DPC αi and IC βi , respectively, and the ranges of DPC and IC are equally set to all nodes. Note that αmax is set in accordance with a given lower limit compression ratio per node Rtrg (see also Section III-A), in order to explicitly control the amount of data transfer traffic. The parameters and functions used in our formulations are summarized in Table I.
Atrg Rtrg Ce Cd Cs E S Function Acc() Cost() Edd () DD() or DD() Size()
The list of original datasets: O = [O1 , O2 , · · · , ON ] (Oi : The original dataset at the ith edge) The list of distilled datasets: D = [D1 , D2 , · · · , DN ] (Di : The distilled dataset at the ith edge) The parameter to setup the target accuracy The parameter to setup the lower limit compression ratio per node e The vector of edge energy prices [$/J]: Ce = [C1e , · · · , CN ] d The vector of data transfer prices [$/B]: Cd = [C1d , · · · , CN ] The cloud storage cost per size (e.g., $ 0.023 per GB [83]) The vector of DD energy overhead: E = [E1 , · · · , EN ] The vector of data size to transfer: S = [S1 , · · · , SN ] Remarks The test accuracy for the aggregated distilled dataset D = [D1 , · · · , DN ] The total cost as a function of O, α, β, Ce , Cd , Cs DD’s energy overhead (Oi → Di ) as a function of Oi , αi , βi Dataset distillation (Oi → Di ) or (O → D) as a function of (Oi αi , βi ) or (O, α, β) The size of the given dataset Di or set of dataset D
energy consumption of the main loop can be considered a linear function of βi , the iteration count. Here, the intercept d4,i can be interpreted as the initialization overhead before the main loop. As for the parameter αi , we approximate its impact on per-loop runtime or energy using the quadratic function (see also Section III-C), which is validated later in Section VI-B1. For the size function, we apply the simple analytical model described in Section III-A. More specifically, the size of a distilled dataset at the ith edge (Di ) is formulated as follows: Si = Size(Di = DD(Oi , αi , βi )) = KV αi
Online Phase
Offline Phase: Energy Model Calibration
Dataset DB O
Model DB {d1,i, ..., d4,i}
Acccuracy Model Calibration (Skipped When Applying a Model from DB) Acc()
ααii,, ββii ααα,,ii,,βββiβi ii Edge ααiα,,iiβ βi
Curve Curve Edge DD j j Curve Edge DD Node Fitting Edge {α,i,ββ, i,EE}i}Curve DD Node Fitting {α DD Node i β ,i E }i Fitting , {α Node {α i i i Fitting Eij} i ji, βEdge ji,Edge Edgei i Edge i
DD
Learning & Fitting
Model DB
CEDD-calibrator
DD Parameter Optimization
C e, C d , C s
Edd() Acc() or Acc()
Edges
CEDD-calibrator
Cost DB
α, β Solver
DD
Training Host
Edges
CEDD-solver
Fig. 9. Overview of CEDD-optimizer Workflow
(8)
Here, K is the number of classes, while V is the size of each data in the dataset. Note, DD() represents the DD procedure at ith node, i.e., returning the distilled dataset Di for a given input set of Oi , αi , and βi . 2) Statistical Accuracy Modeling: To constrain training quality, we model the expected test accuracy of the downstream deep learning task trained with the aggregated set of distilled datasets (D = DD(O, α, β)). While several early studies have also attempted to model the test accuracy [84], [85], our work newly considers the variation of hyperparameter configurations uniquely set across N devices in the modeling. To this end, we treat the variation of hyperparameters α, β statistically and macroscopically for the following reasons: • The test accuracy is permutation invariant with respect to the elements in each hyperparameter vector α / β particularly for IID data (see Section VII-F for non-IID cases). • For a permutation-invariant function f (x), there exist funcP tions ρ and ϕ that meet f (x) = ρ( x∈x ϕ(x)) [86]. • Any sample moment can be expressed with the form of P x∈x ϕ(x), and the first, second, third, fourth, · · · order sample moments correspond to descriptive statistics, namely mean, variance, skewness, kurtosis, · · · [87]. • After eliminating irrelevant statistics, a polynomial function of the remaining descriptive statistics should well represent the test accuracy in manner of Maclaurin expansion. In this work, we stop at the second order moment, and apply the following second order polynomial form by using means (µα , µβ ) and standard deviations (σα , σβ ) of α and β:
V. S YSTEM D ESIGN A. Solution Overview
Fig. 9 illustrates the overall workflow of CEDD-optimizer to solve the optimization problem presented in Section IV-A by using the model introduced in Section IV-B. As shown, our solution consists of two modules: (1) the CEDD-calibrator and (2) the CEDD-solver. The CEDD-calibrator deals with online or offline model parameter identification, while the CEDD-solver optimizes the hyperparameter setup online. In the offline phase, the CEDD-calibrator runs the DD code on a benchmark dataset while varying the hyperparameters α and β, measures the energy consumption per iteration, and fits the polynomial curve described by Eq. (7). Note that this model calibration procedure is required only once per edge device. The calibration comprises two steps: measuring d4,i (initialization and finalization overhead) by setting β = 0; and exploring α within [αmin , αmax ], while fixing β at 1. In the online phase, both modules are involved. The CEDDsolver solves the optimization problem by using the price parameters (Ce , Cd , Cs ) stored in the cost database. In our implementation, we use scipy.optimize in SciPy with SLSQP option to solve the problem [88]. Our solver requires the coefficients of the energy and test accuracy models as inputs. For the former, we use the coefficients identified in the offline calibration mentioned above, while for the latter, we offer two options to cope with the dependency of test accuracy on dataset features. The default option is reactive, agnostic to the dataset — it learns the model coefficients online using 2 a three-step tuning algorithm, described later in Section V-B. Acc(D) ≃ a0 + a1 µα + a2 µβ + a3 µα + a4 µα µβ After this online model calibration is completed, the learned 2 2 2 +a5 µβ + a6 σα + a7 σβ (9) coefficients are stored in the model database for record or This simple form of the test accuracy model is beneficial in potential future reuse. Another option is proactive, using prior limiting the calibration overhead and scaling the system because knowledge provided by the user — it applies a user defined or the number of unknowns is constant as the edge count N scales. pre-trained test accuracy model stored in the model database At the same time, limiting the orders of moment and polynomial (e.g., reuse a model from a previous run). The approach is contributes to avoiding overfitting, particularly compared to a valid as long as the user is confident with the coefficient setup model that fully uses all parameters in α and β with a high and reduces the overhead by skipping the online training. polynomial order. Note that one can improve the accuracy of this function simply (1) by using additional descriptive B. Three-Step Hyperparameter Tuning Scheme statistics (e.g., skewness and kurtosis) or (2) by increasing the Fig. 10 illustrates the online three-step tuning approach that order of the polynomial approximation. Nevertheless, we use coordinates the CEDD-calibrator and the CEDD-solver. Overall, the above simple form based on our preliminary experiments, the scheme aims to efficiently generate a training dataset that which is also validated later in Section VI-B13. Consequently, fits the test accuracy curve. it is lightweight but accurate enough for various edge counts In the first step, DD is applied across all edges while N as demonstrated later in Section VI-B. setting α and β to the minimum, i.e., αmin = 1 and
EdgeN
Edge1 Step1: DD w/ Minimum Setup
β
Step2: DD w/ Random Setup
β
・・・
・ Δβ
α
・・・
・
Minimum Setup
β
・
α
DD DD DD DD α&β Random Setup
α
α&β
β
・
Δα α
DD DD DD DD
Dt1, Dt2, . . .
Downstream NN Model Training
Universe μβ Extrapolation Accuracy
・ ・・ ・ ・・ ・・
Set of Datasets Dr = [Dr1, . . ., DrN]
Calibration Datasets Creation Dtj = [Dtj1, …, DtjN], Dtji = Dmi or Dri
Accuracy
Online Calibration
Set of Datasets Dm = [Dm1, . . ., DmN]
Universe
σβ Accuracy μα
・ ・・・ ・・・ ・
Dtj
・・・ σα
Basic Statistics of α & β Sets of Accuracy & Basic Statistics
Finally, CEDD-solver optimizes the setups of α and β by using the test accuracy model, the energy model, and the price parameters, then DD is performed across edge devices, with the optimal hyperparameter setting. Our scheme takes three or less steps to find the optimal setting, while limiting the overhead spent on the first two steps. More specifically, ours can stop at the end of the first or second step if the test accuracy meets the target Atrg for Dm or Dr . VI. E VALUATION
Curve Fitting Accuracy Curve Step3: DD w/ Optimal Setup
Optimal Setup DD DD DD DD
CEDD-solver
Set of Datasets D = [D1, . . ., DN]
α&β
Fig. 10. Online Three-Step Hyperparameter Tuning
βmin = 1, respectively. Then, a set of distilled datasets m m Dm = [Dm 1 , · · · , DN ] is generated where Di represents the dataset distilled at the ith edge node. As the energy consumption of DD is modeled as O(α2 β), as denoted in Eq. (7), the overhead of this step is almost negligible. The second step then creates N sets of (α, β) where α and β are randomly chosen from [αmin , αmin + ∆α ] and [βmin , βmin + ∆β ], respectively. Note that the total cost, including the DD energy consumption and the subsequent data transfer, is controlled by ∆α and ∆β . Once the set of (α, β) is generated, its mapping to N edge devices is optimized to minimize the total cost, encompassing both energy consumption and data transfer across N edge devices, by solving the minimum-weight perfect bipartite matching problem (MWPBM). Then, an assignment solver namely linear_sum_assignment() [89] in SciPy is used, which solves the problem using the Jonker-Volgenant algorithm [90]. Consequently, a set of distilled datasets Dr = [Dr1 , · · · , DrN ] is collected (Dri : the distilled dataset at the ith edge node). Once the datasets Dm and Dr are collected at the host node, multipleh training datasets are created Dt1 , Dt2 , · · · where i t t t r Dtj = D1j , · · · , DNj (Di j = Dm i or Di , i: edge ID). Here, the jth training dataset Dtj is composed by choosing the distilled dataset generated either in the first or second step r (Dm i or Di ) for each edge i, which is randomly selected. The set of α and β setups across N edges is associated with Dtj , and the basic statistics (e.g., µα ) are calculated. Then, the training host performs the downstream neural network training and outputs the test accuracy for all training datasets Dt1 , Dt2 , · · · . As these datasets are distilled and, hence, very compact, the overhead of this training is very limited as demonstrated later in Section VI-B (4th graph in Fig 15). We then combine the test accuracy output and the basic statistics for each training dataset Dtj . We fit the test accuracy curve modeled in Section IV-B2 to the combinations of test accuracy and basic statistics across all training datasets for interpolation and extrapolation. To guide this fitting procedure, predefined boundary information is used. Note that the accuracy of this test accuracy curve is controllable by ∆α and ∆β .
A. Evaluation Setup This section describes the details of our evaluation setup. The default parameter setup is listed in Table II, but we adjust them for each experiment based on the experiment purpose, as documented in the relevant sections. 1) Hardware Platforms: We focus on the following devices: (1) NVIDIA Jetson Xavier NX; (2) NVIDIA Jetson Orin Nano; (3) MacBook Pro; and (4) TQ module TQMx80UC. The first two are embedded computing platforms designed for edge and AI computing. The Xavier NX comprises a 6core ARM Carmel CPU and a 384-core Volta GPU, while the Orin Nano is composed of 6-core ARM Cortex-A78AE CPU and an Ampere GPU. Both are representatives of edge-AI computing systems, powered by Tensor Cores that enable energy-efficient training and inferences for deep learning tasks. The MacBook Pro Quad-Core is a powerful mobile computing platform, consisting of an Intel Core-i5 processor and a discrete GPU. The TQMx80UC COM module [91] is an x86-based edge computing platform, comprising a long-lifecycle Intel core processor (8th generation). These systems cover a wide spectrum of edge and mobile computing platforms, are suitable for our target scenario, and ensure the robustness of our experiments with respect to hardware characteristics. 2) Baseline Datasets and Networks: We use publicly available standard datasets to illustrate the problem in a controlled and reproducible manner: MNIST [77], FashionMNIST [92], CIFAR-10 [78], SVHN [93], ImageNette [94] (subset images of ImageNet [95]). The applicability to even larger datasets is discussed in Section VI-B8. While we use these image datasets following the literature, very recent studies report that DD is also applicable to time series data [96]. We evaluate our methodology under two scaling regimes, in terms of the number of edge devices N : Strong Scaling, i.e., the total dataset size remains fixed while the number of edge devices increases; and Weak Scaling, i.e., the dataset size per edge node remains constant when the number of edge devices scales. We focus on three data-distribution settings:(1) Independent and Identically Distributed (IID) data; (2) quantity-based
TABLE II D EFAULT PARAMETER S ETUP FOR DD Parameter Setup αmin =βmin =1, αmax =70, βmax =40, ∆α=25, ∆β =15, N =10, Rtrg =M/(N Kαmax ), dataset = MNIST [77], host = MacBook, edges = TQMx80UC x2, MacBook x3, Jeston Orin x2, Jetson Xavier x3
near-IID data generated using Dirichlet distribution with concentration parameter of α = 1.0; (2) quantity-based non-IID data generated using Dirichlet distribution with concentration parameter of α = 0.5. In all settings, we ensure that each edge device contains at least 15 samples from each class. For our test accuracy model, the IID setup is more suitable by its design. A potential extension is discussed later in Section VII-F. For our downstream deep learning tasks, we use a neural network architecture designed by Gidaris and Komodakis [97]. The architecture encompasses (1) three convolutional blocks with 3x3 filters, (2) instance normalization [98], (3) RELU activation, and (4) 2×2 average pooling with a stride of 2. 3) DD Implementation: Among the DD methods introduced in Section II-A, we select the vanilla Dataset Condensation (DC) [12] as our DD baseline. This is because it exhibits fundamental structures commonly found in most other mainstream studies and implementations [14], [15], [29], [30], [38], [99]. As we apply a parametrization to the commonly available basic structure, our approach is broadly applicable to other DD frameworks that may integrate sophisticated approaches/concepts, such as improved network architectures [39], [100], label learning [40]–[42], long-range trajectories [15], [30], and others [14], [29], [101]–[103]. 4) Energy Model Calibration: To identify the coefficients of our energy model in Eq. (7), we measure the average runtime and power while scaling DPC (α) selected from {1, 10, 20, 30, 40, 50, 60, 70}. We measure power consumption using the following tools: tegrastats for Jetson Xavier NX and Jetson Orin Nano; powermetrics for MacBook Pro; and powerstat for the TQMx80UC. Note that the clock frequency is set to the default for all, and coordination with clock scaling in our approach is orthogonal to our optimization, as discussed later in Section VII. For the energy calibration, we divide the DD execution into two phases as mentioned in Section V-A: (a) the initialization (i.e., encompassing module loading and data initialization), and (b) the main loop of the DD iterative solver. 5) Price Parameter Setup: The energy and size models are then used to quantify the total cost formulated in Eq. (5). To this end, the price parameters Ce , Cd , and Cs need to be identified beforehand and depend heavily on the geographical locations of edge servers, as previously mentioned. For the energy price parameters Ce , we use the same data presented in Fig. 6, which is based on the World Population Review [80]1 . As for the network price parameters Cd , we extract the operating cost per region (SiteLink Rate) from [81], which is the same as Fig. 7. Regarding the storage price parameter, we model it after Amazon’s S3 pricing [83] and set it to $ 0.023 per GB, regardless of the edge locations. As for the geographical location setting, we first choose 74 out of the 144 countries listed in the World Population Review [80] based on the applicability of the list shown in Fig. 7. We then set the geographical location for each of the N edge devices by randomly choosing one of the 74 countries. 1 Although the electricity price depends also on the contract and varies within a country, the parameter setups still covers the real-world prices.
6) Baseline Methods: We compare against a DD baseline with (α, β) = (αmax , βmax ) and conventional random sampling [12], [29], [30]. We further compare ours with commonly-used surrogate-free iterative search algorithms (i.e., random search, heuristic search, exhaustive search [104]). The advantage of DD over generic lossy/lossless compression methods is shown in Fig. 3. Note, the compression methods either do not satisfy the target compression ratio Rtrg or significantly violate the target accuracy Atrg . Consequently, we do not include them in our comparisons. B. Experimental Result 1) Validation of the Energy Model: We perform our energy model experiments on the following platforms: MacBook Pro, TQMx80UC, Jeston Xiavier NX, and Jeston Orin Nano using the setup detailed in Section VI-A4. Fig. 11 visualizes the ground truth, fitted surface, and predicted values for these four platforms as well as the comparison of energy surfaces among the platforms. Overall, the energy model matches the ground truth accurately, as the predicted values are very close to the ground truth energy. The prediction errors are generally very small for all devices — lower than 3%. As shown in the rightmost figure, the shapes of energy surfaces differ significantly across devices, underscoring the importance of hardware-aware hyperparameter tuning when the system is heterogeneous in terms of edge hardware. Our tuning approach covers this aspect by simply using device-specific coefficients (d1,i , · · · , d4,i ) for each edge node. 2) Validation of Accuracy Model: We estimate the test-accuracy model coefficients as follows. First, we randomly generate 18 sets of α = [α1 , α2 , · · · , αN ] and β = [β1 , β2 , · · · , βN ]. For each hyperparameter setup, we run DD on every edge server, aggregate the distilled data centrally, and train/evaluate a randomly initialized network on the aggregated set. This yields 18 test accuracies Acc, each averaged over five variants, together with the corresponding descriptive statistics (µα , · · · ). Note, we found that 18 samples are sufficient for stable performance based on a pilot study with up to 100 samples. We split the data into 8 training and 10 test points for fitting Eq. (9) using stochastic gradient descent (SGD) [105], [106], yielding coefficients a0 , a1 , . . . , a7 . We present the errors of our test accuracy model (Acc()) using various datasets while scaling the number of edge nodes N by using the setup detailed in Section IV-B2. Fig. 12 shows the result for both weak and strong scaling (defined in Section VI-A2): (a) absolute error and (b) Mean Absolute Percentage Error (MAPE). For MNIST [77], FashionMNIST [92], SVHN [93] and CIFAR-10 [78], the absolute error and MAPE are smaller than 0.5 and 1%, respectively, when the number of edges N is greater than 10. Generally, the error becomes smaller as we scale N , and saturates at a certain point. We assume this behavior is caused by its statistical nature (i.e., the central limit theorem) as we utilize several descriptive statistics such as the means in the test accuracy model (Acc()). 3) Sample Count Selection in CEDD-calibrator: Fig. 13 presents the error as a function of the number of sampling
0 10
20
30
0
600 400 200 0 30 20 10
0 10
20
30
0
400 200 0 30 20 10
100
0 10
20
30
0
0 30 20 10
0 10
20
30
Energy Comparison
0
400 200 0 30 20 10
MacBook Pro TQMx80UC Jeston Xiavier NX Jeston Orin Nano
0 10
20
30
Energy(kJ)
Prediction surface Real samples Predicted samples
Energy(kJ)
Jeston Orin Nano
Prediction surface Real samples Predicted samples
Energy(kJ)
Jeston Xiavier NX Energy(kJ)
TQMx80UC Prediction surface Real samples Predicted samples
Energy(kJ)
MacBook Pro Prediction surface Real samples Predicted samples
600 400 200
0
30 20 10
Fig. 11. Time and Power Measurement and Their Approximation on Various Devices
(a) Accuracy Prediction Error 2.0 1.5 1.0 0.5 0.0
0
20
40
N
60
80
MNIST, weak SVHN, weak CIFAR-10, weak FasionMNIST, weak ImageNette, weak MNIST, strong SVHN, strong CIFAR-10, strong FasionMNIST, strong ImageNette, strong
5 4 MAPE %
Absolute Error
(b) Accuracy Prediction Error
MNIST, weak SVHN, weak CIFAR-10, weak FasionMNIST, weak ImageNette, weak MNIST, strong SVHN, strong CIFAR-10, strong FasionMNIST, strong ImageNette, strong
2.5
3 2 1 0
100
0
20
40
N
60
80
100
Fig. 12. Error v.s. Scale in Our Test Accuracy Model Acc() Energy Prediction Elbow Point 25 MAPE(%)
20 15 Elbow point
10
Accuracy Prediction Elbow Point
MacBook Pro MNIST MacBook Pro Cifar10 MacBook Pro FasionMNIST MacBook Pro SVHN MacBook Pro ImageNette TQMx80UC MNIST TQMx80UC Cifar10 TQMx80UC FasionMNIST TQMx80UC SVHN TQMx80UC ImageNette
40 30 20 Elbow point
10
5 0
CIFAR-10 SVHN FashionMNIST MNIST ImageNette
50
MAPE(%)
30
0 1
2
3 4 5 Number of Data Samples
6
7
2
4
6 8 10 12 Number of Data Samples
14
Fig. 13. Model Error as a Function of Fitting Sample Count
points to fit for our energy and test accuracy modeling. As an error metric, we again use MAPE. The X-axis indicates the number of samples, while the Y-axis represents MAPE [%]. Our accuracy and energy models achieve MAPEs below 6% and 3%, respectively, across various datasets. Note, an MAPE below 10% is generally considered a good fit. Based on this experiment, we judge that 8 and 3 fitting points are sufficient for our accuracy or energy models, respectively. In our CEDDCalibrator implementation, we sample more to set margins and test the curves after the fitting. Consequently, we sample 18 or 5 points for the accuracy or energy models, respectively. 4) Cost and Test Accuracy Benefits by CEDD-optimizer: Fig. 14 showcases the effectiveness of our CEDD-optimizer across 5 standard datasets. The horizontal axis indicates the accuracy, whereas the vertical axis indicates the total cost encompassing DD energy consumption, data transfer, cloud storage, and host energy consumption. Intuitively, the lower right region is better on the 2D plane. In this experiment, we compare the following methods in terms of cost and accuracy: (a) One Step, finishing the tuning after the first step in CEDD-optimizer (α = β = 1); (b) Two Step, finishing the
tuning after the second step; (c) Three Step, performing the whole procedures in our three step tuning; (d) Random Sampling, randomly choosing a subset of the original dataset in each edge device such that the total data transfer size becomes the same as that of Three Step (the same size for all edges); (e) DD Baseline, applying DD at α = αmax and β = βmax ; and (f) Ideal, CEDD-optimizer when using the accuracy model from the database that is trained with ∆α = αmax − αmin and ∆β = βmax − βmin (ideal scenario). Overall, our three-step tuning scheme significantly reduces the cost compared with the DD Baseline, while keeping almost the same accuracy. As implied by the cost difference between Three Step and Ideal, one can achieve an even further cost reduction by reusing the test accuracy model, if applicable. Compared with Random Sampling, Ideal improves accuracy significantly while keeping the cost almost the same, implying that tuning is as simple as Random Sampling. CEDD-ideal reduces cost by 5.25×– 20.8× and the three-step algorithm by 1.7×–2.67× versus the DD baseline, both with only a 1%–1.5% accuracy drop. 5) Impact of ∆α and ∆β on Accuracy Model Quality in Our Three Step Tuning: The effectiveness of our three-step tuning scheme depends on the setup of ∆α and ∆β . The lower ∆α and ∆β are set, the less cost the calibration requires, but also the less accurate the test accuracy model becomes, as these parameters control the ranges of α and β selected in the second step (see Section V-B. As ∆α and ∆β are set smaller, the covered regions of basic statistics (e.g., µα ) also become smaller, which degrades the test accuracy function’s accuracy as a consequence. The leftmost heatmap in Fig. 15 visualizes the impact of ∆α and ∆β on the accuracy of the test accuracy model. The gray region (∆α = ∆β = 5) indicates a failure case, in which the solver does not find a solution. Based on the observation, we set (∆α , ∆β ) = (25, 15), which well balances the trade-off throughout our experiments. 6) Sensitivity to Device/Location Heterogeneity: The plots numbered (2) and (3) in Fig. 15 demonstrate our two ablation studies conducted for MNIST, i.e., removing the heterogeneity in terms of edge hardware and location from the default. More specifically, we use the same edge device, namely Jetson Xavier, for the former, while the device location is set uniformly for the latter. The other evaluation parameters are unchanged from the default. Overall, our three-step tuning scheme is still effective compared with
0.5
3
0.5
3
Cost($)
1.0 Cost($)
Cost($)
1.0
MNIST
2 2 0.0 1 40 60 60 80 Accuracy(%) Accuracy(%) One Step Two Steps Three Steps
1.0 0.5
3
1.0 0.5
3
0.0
ImageNette
15
2 2 0.0 1 95 100 80 90 Accuracy(%) Accuracy(%) Random Sampling DD Baseline Ideal
1
0.0
Fashion-MNIST
10
Cost($)
SVHN
Cost($)
CIFAR-10
3
5
1
0
2
1
25 50 75 Accuracy(%) Step Progression (1 2 3)
Fig. 14. Cost and Test Accuracy Comparison across Various Standard Datasets
5 5
15 25 35 45
97
0.6
96
0.4
95
0.2
94
0.0
(3) Uniform Location One Step Two Steps Three Steps Random Sampling DD Baseline
1.5 1.0
3
0.5
1
92.5
2
95.0 97.5 100.0 Accuracy(%)
0.0
1
92.5
(4) Training 5 Epochs 6000
2 95.0 97.5 100.0 Accuracy(%)
4000 2000 0
(5) Solver Runtime(s)
6098.8
log10(Runtime)
15
0.8
2.0
Energy(J)
25
98
One Step Two Steps Three Steps Random Sampling DD Baseline 3
Cost($)
35
(2) Uniform Device
99 Accuracy(%) Cost($)
(1) Accuracy over ( , ) 45
19.3 115.6 308.2
=1) 10) 25) rain DD( DD( = DD( =tandard T S
2 1 0 1
1.0
1.5 2.0 2.5 log10(EdgeCount)
Fig. 15. A Compilation of Experimental Results: (1) Impact of ∆α and ∆β on Accuracy of Test Accuracy Model; (2)/(3) Test Accuracy v.s. Cost for Uniform Device/Location Setup; (4) Comparison of Energy Overhead for Downstream Training Task; and (5) Solver Runtime as a Function of Edge Count N
the DD Baseline and Random Sampling for these less a function of N , thus the total energy cost for DD is O(N ) for heterogeneous environments. both scaling policies. Moreover, as the volume of the distilled 7) Cost Reduction by DD for Downstream Learning: Our dataset per node is denoted as KV α (see Section IV-B1), three-step tuning scheme requires multiple downstream training the data transfer and storage costs are also O(N ) for both runs to generate samples for fitting the test-accuracy curve. scaling policies. As the curve to fit in our approach is scale However, this overhead is limited because downstream training free (see Section IV-B2), the fitting overhead is O(1) for both is performed on distilled datasets. Fig. 15 (4) shows this effect policies. As for the cost of the downstream learning, Assuming using datasets distilled under different hyperparameter setups the number of training epochs scales proportionally with the The vertical axis reports the training energy consumption over total dataset volume while the batch size remains constant, the five epochs, while the horizontal axis lists the downstream learning overhead scales as O(N ) under both scaling policies. training with different datasets. Specifically, DD(α=X) shows However, the overhead of the solver for tuning hyperparameters dataset distilled with α = X, whereas Standard Train is O(N 3 ), and thus this can cause bottlenecks at large scale. utilizes the original dataset without DD with batch size equals In Fig. 15 (5) represents the solver’s run time as a function of 512. Here, experiments are on MNIST. Since this significant edge device count N . In our environment, the solver takes 3.4s training cost reduction in the downstream training task, Our at N = 100, but it can become a bottleneck at larger scale. One three-step tuning scheme is effective as shown in Fig. 14. prominent solution is simply to replace the algorithm (SLSQP) 8) Scalability Assessment: In this section, we assess the with a more efficient heuristic alternative. scalability of our CEDD-optimizer in terms of cost by analyzing Under weak scaling, the total cost (excluding the solver) the scaling properties. Note that the scaling rule of the test scales as O(N ) both with and without DD. Consequently, accuracy error is presented and discussed in Section VI-B2. the cost-reduction ratio relative to the “without DD” baseline Here, we consider two scaling policies: weak and strong scaling. remains O(1), i.e., DD provides scale-invariant benefits under For the former, we assume that the total data count M across weak scaling. Under strong scaling, the cost-reduction ratio all edge devices is scaled proportionally to the edge count N , while the number of classes K, the size per data V , and the distributions of α and β are kept constant. As for the latter, TABLE III we assume that the total data count M remains constant as the C OST S CALING RULES IN CEDD- OPTIMIZER edge device count N scales. The other scaling conditions for Total DD Traffic & Scaling Policy Fitting Learning Solver K, V , α, and β are the same as those of weak scaling. Energy Storage Table III lists scaling factors for various costs in both scaling Weak w/ DD O(N ) O(N ) O(1) O(N ) O(N 3 ) policies with or without DD. The scaling factors of ”with DD” Weak w/o DD – O(N ) – O(N ) – Strong w/ DD O(N ) O(N ) O(1) O(N ) O(N 3 ) are derived as follows: First, as mentioned in Section IV-B1, Strong w/o DD – O(1) – O(1) – the energy overhead of DD per edge is O(α2 β), which is not
Fashion-MNIST (IID) CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
2
0 40
60 80 Accuracy(%)
2 0
80
90 Accuracy(%)
100
80 Accuracy(%)
SVHN (Near IID)
MNIST (Near IID)
Fashion-MNIST (Near IID)
CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
4 2
70
4 2
6
0
0 40 50 60 Accuracy(%)
6
40
60 80 Accuracy(%)
4 2
90 Accuracy(%)
100
40
50 60 Accuracy(%)
70
ImageNette (Near IID) CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
100
0 80
50
90
CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
6
CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
0 70
CIFAR-10 (Near IID)
0 30
2
4
Cost($)
4
70
Cost($)
6
40 50 60 Accuracy(%)
Cost($)
30
4
6
ImageNette (IID) 100 Cost($)
2 0
0
Cost($)
4
6
Cost($)
6
Cost($)
2
MNIST (IID) CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
Cost($)
4
SVHN (IID) CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
Cost($)
Cost($)
6
CIFAR-10 (IID) CEDD-optimizer Random Sampling Ideal DD Baseline Exhaustive Search Random Search Heuristic Search
50 0
70
80 Accuracy(%)
90
40
50 60 Accuracy(%)
70
Fig. 16. Accuracy-Cost Comparisons with Baselines and Iterative Search Methods under IID and Near-IID Scenarios
decreases with N by a factor of O(1/N ). In practice, we expect the possible uniform α and β settings, draw the Pareto frontier weak scaling to be more common, since deploying additional of the uniform setups on the accuracy-cost plane, and compare edges is typically driven by the need to collect more data. our non-uniform approach with it. Note that the same identical 9) Comparison to Iterative Search Methods: It is hard set of α and β is applied to all edge nodes for the uniform for Random sampling to reach a near Ideal perfor- setups, while α and β can be different values. The four leftmost plots in Fig. 17 compare our approach with mance, while merely using DD baseline is also inefficient. Our three-step approach efficiently searches an optimal con- the uniform setups and their Pareto frontier for CIFAR-10 figuration, while intelligently limiting the search overhead. and SVHN under IID and non-IID scenarios. Across all four Fig. 16 showcases its efficiency by comparing it with three dataset distribution scenarios, our approach is better than the commonly-used iterative search algorithms and the baselines. Pareto frontier. In practice, the benefit would be even more Exhaustive search is implemented as a grid search considerable as it would not always be possible to select a over (α ∈ 1, 15, 35, 55, 70) and (β ∈ 1, 10, 15, 20, 30), while Pareto optimal setup, and suboptimal uniform setup points are Random search iteratively evaluates 20 randomly generated mostly far from our optimizer on this plane. (α, β) configurations. Heuristic search starts from the 11) Impact of Network Depth on Cost Saving: The smallest parameter values and increases α in steps of 15 and β rightmost plot in Fig. 17 evaluates how the network depth in steps of 10 until the accuracy constraint is satisfied. These or size setup for DD affects the benefit of our method three methods iteratively search an optimal hypearparameter in the previously mentioned Ideal case. We evaluate this configuration in a step-by-step manner. Each iteration requires impact under the assumption that the test accuracy function both the DD and the downstream training task, and the overhead remains unchanged for different network depths, in order to depends on the choice of the hyperparameter vectors α and β. eliminate cost change factors other than the network depth, most Exhaustive search and Heuristic search gen- prominently the hyperparameter setups α and β. The benefit of erally satisfy the constraint under both IID and near-IID data, our method gradually increases as the network becomes deeper: whereas Random search can fail to find a feasible setup the cost ratio of DD baseline to ours increases modestly within the 20 attempts. Nevertheless, these three approaches are from approximately 20.8× for three layers to slightly over even more expensive than DD baseline that performs only 22× for four and five layers. one DD run with the maximum parameter configuration (αi , 12) Impact of Solver Selection in CEDD-optimizer: We βi )=(αmax , βmax ). CEDD-optimizer intelligently avoids then evaluate the impact of solver selection in the optimizer this situation by (1) restricting the search space with ∆α on overhead and quality, namely, the solver run time and the and ∆β to control the overhead per search, (2) limiting the objective value. To this end, we compare the SLSQP solver with number of search iterations while fitting a simple curve, and some alternatives: (1) a Trust-Region Constrained optimization (3) pinpointing an optimal with a solver after the fitting. (TRC) solver [107], [108]; (2) an Integer Differential Evolution 10) Benefit of Non-Uniform Hyperparameter Setup: Our (IDE) solver [109], [110]; and (3) a Mixed-Integer Nonlinear CEDD-optimizer aims to configure the DD hyperparameters Programming (MINLP) solver [111], [112]. Note that the α and β uniquely across edge devices depending on the SLSQP and TRC solvers are continuous, while the rest are heterogeneous price parameters. We then assess the benefit discrete solvers. Therefore, continuous relaxation is applied to of this non-uniform tuning compared with uniform setups, in the former two. terms of the cost and the test accuracy. To do so, we test all The two leftmost graphs in Fig. 18 compare the solvers in
Network Depth Impact Cost Reduction Ratio
Pareto Frontier (SVHN Non-IID) Pareto Frontier (SVHN IID) Uniform Configs Uniform Configs 1.25 1.25 Pareto Optimal Pareto Optimal Pareto Frontier Pareto Frontier 1.00 1.00 Our Optimizer Our Optimizer 0.75 0.75 0.50 0.50 0.25 0.25 0.00 0.00 96 97 98 96 98 Accuracy(%) Accuracy(%) Cost($)
Cost($)
Cost($)
Cost($)
Pareto Frontier (CIFAR-10 Non-IID) Pareto Frontier (CIFAR-10 IID) Uniform Configs Uniform Configs 1.25 1.25 Pareto Optimal Pareto Optimal Pareto Frontier Pareto Frontier 1.00 1.00 Our Optimizer Our Optimizer 0.75 0.75 0.50 0.50 0.25 0.25 0.00 0.00 55 60 65 50 60 Accuracy(%) Accuracy(%)
20 15 10 5 0
3 4 5 Number of Layers
Fig. 17. Our Optimizer with a Non-Uniform Hyperparameter Setup v.s. Uniform Hyperparameter Setups under IID and Non-IID Scenarios (Four Leftmost Graphs) and Impact of Network Depth on Cost Saving Factor (Rightmost)
0 1 0.006 s
SLSQP TRC (Ours)
IDE
MINLP
1.4 1.2 1.0 0.8 0.6 0.4 0.2 0.0
1.33× 1.00×
(c) Absolute Error
System Cost Optimizer Cost
0.94× 0.91×
2.5 2.0
0.5 SLSQP TRC (Ours)
IDE
MINLP
0.0
4.65
4 1.75
1.5 1.0
(d) MAPE 5
2.59
MAPE (%)
2.623 s
2.431 s
Relative Cost
log10(Time (s))
1
2
(b) Relative Total Cost
110.686 s
Absolute Error
(a) Optimizer Time 2
0.78 0.730.720.72
0.64 0.57
A B C D E F G H Accuracy Model
3 2 1 0
3.12
1.40 1.311.301.29
1.17 1.03
A B C D E F G H Accuracy Model
Fig. 18. Time (a) and Cost (b) Comparison with Alternative Solvers and Test Accuracy Model Comparison in Absolute Error (c) and MAPE (d)
terms of run time and total cost, respectively. In this comparison, we set the edge count N to 10. As shown in the graph (a), the run time is only 0.006s for the SLSQP solver, whereas it takes 2.431s, 110.686s, and 2.623s for the TRC, IDE, and MINLP solvers, respectively. The graph (b) reports the total cost for the aforementioned Ideal scenario. The total cost is broken down into (1) system cost and (2) optimizer (or solver) cost, which are normalized to the total cost of SLSQP. The relative costs of TRC, IDE, and MINLP are 1.33, 0.94, and 0.91, respectively. Consequently, the SLSQP solver outputs a solution comparable to those of the discrete solvers within orders of magnitude shorter run time. Furthermore, these discrete solvers do not solve the problem in a practical amount of time when N is scaled, due to the complex nature of the integer optimization. The SLSQP solver, therefore, is a favorable choice to balance solution quality, overhead, and scalability. 13) Order Selection in Test Accuracy Model: Next, we validate the choice of polynomial degree and moment order in our test accuracy function formulated in Section IV-B2. To this end, we define the variables of skewness and kurtosis as follows: sα = skew(α), sβ = skew(β), kα = kurt(α), kβ = kurt(β). We compare the following 8 different models to represent the test accuracy function Acc().
a4 µα µβ + a5 µ2β + a6 σα2 + a7 σβ2 . E. Second-order means and variances (11 unknowns): a0 + a1 µα + a2 µβ + a3 µ2α + a4 µα µβ + a5 µ2β + a6 σα2 + a7 σβ2 + a8 σα4 + a9 σα2 σβ2 + a10 σβ4 . F. Third-order means and second-order variances (15 unknowns): a0 + a1 µα + a2 µβ + a3 µ2α + a4 µα µβ + a5 µ2β +a6 σα2 +a7 σβ2 +a8 σα4 +a9 σα2 σβ2 +a10 σβ4 +a11 µ3α + a12 µ2α µβ + a13 µα µ2β + a14 µ3β . G. Model D (our choice) plus first- and second-order skewness (13 unknowns): a0 + a1 µα + a2 µβ + a3 µ2α + a4 µα µβ + a5 µ2β + a6 σα2 + a7 σβ2 + b1 sα + b2 sβ + b3 s2α + b4 sα sβ + b5 s2β . H. Model G plus first- and second-order kurtosis (18 unknowns): a0 +a1 µα +a2 µβ +a3 µ2α +a4 µα µβ +a5 µ2β + a6 σα2 + a7 σβ2 + b1 sα + b2 sβ + b3 s2α + b4 sα sβ + b5 s2β + b6 kα + b7 kβ + b8 kα2 + b9 kα kβ + b10 kβ2 .
Fig. 18 (c) / (d) compares the above functions in absolute error / MAPE. The model D (our choice) achieves an absolute error of 0.73 and an MAPE of 1.31%, closely matching those of the models E and F. This means that increasing the degree of polynomial of our model does not contribute significantly to error reduction. As for the order of statistical moment, scaling A. First-order means (3 unknowns): a0 + a1 µα + a2 µβ . the order by adding more descriptive statistics gradually reduces B. First-order means and variances (5 unknowns): a0 + the error — see the errors of the models C, D, G, and H. a1 µα + a2 µβ + a6 σα2 + a7 σβ2 . However, a model with more unknowns requires more samples C. Second-order means (6 unknowns): a0 + a1 µα + a2 µβ + in the online calibration process presented in Section V-B. a3 µ2α + a4 µα µβ + a5 µ2β . As an example, the model H requires at least 18 samples. D. Second-order means and first-order variances (our Considering the cost-accuracy trade-off, we stopped at the choice, 8 unknowns): a0 + a1 µα + a2 µβ + a3 µ2α + second order moments by taking only means and variances.
VII. D ISCUSSION In this section, we discuss the impact of DD, as well as additional benefits beyond the results above, and discuss future challenges in its use. A. Privacy Protection As DD strategically distills multiple images into a single image, it deliberately avoids revealing the original data during transmission. Once sensitive information is removed during DD, it is extremely hard to restore it due to its lossy compression nature. We highlight a previous work by Dong et al. [18] that theoretically proves this privacy benefit of DD. Other studies also report the privacy benefit [18]–[20], [102], [113]. Although privacy protection is out of our scope, this preferred feature is another key motivation for and benefit of applying DD.
F. Handling Non-IID Data While CEDD-optimizer mainly targets IID and near-IID data, adapting to general non-IID data would be an immediate next step. First, DD is also valid for non-IID data as reported in various studies [13], [27], [114]–[117]. In our CEDD-optimizer, our cost model is equally applicable to non-IID datasets, while the skewed amount of information across edges, induced by nonIID setting, can affect the test accuracy of downstream training tasks. To take such skewness into account in our test accuracy prediction, a potential solution is simply to use weighted basic statistics (e.g., weighted mean) as the model input (µα , µβ , σα , and σβ ), while keeping the modeling structure unchanged. The key challenge here is quantifying the weights correctly to represent the amount of information each local dataset has. In this calculation, one can use sampled data, instead of using all the data, from the dataset, to limit the overhead.
B. Cross-Architecture Performance The application of DD is agnostic to the target network architecture, which is later trained on the distilled dataset in the downstream training task. This is supported by previous studies demonstrating decent cross-architecture performance [12], [14], [15], [29], [30], [39]. This enables us to adopt a single, simple network architecture for our DD procedure, making it suitable for performing DD on the edges.
G. Temporal Variability of Electricity Prices Electricity prices can vary over time depending on the availability of renewable energy and the power demand. Although we did not take this effect into account in this paper, the variability offers another research opportunity. More specifically, optimizing the trigger timing of DD, using price forecast data, can be another optimization knob that operates on top of our present work.
C. Coordination with Clock Frequency Scaling Another benefit of our approach is the ability to incorporate hyperparameter tuning with power management mechanisms (e.g., clock frequency scaling) across edges. This extension is rather straightforward: simply find the best clock setup during the offline energy-coefficient calibration and apply it in the online DD phase. Consequently, this extension is fundamentally minor and orthogonal to our work, affecting only the energy model coefficients as functions of clock frequency.
VIII. C ONCLUSIONS
This paper explored the use of DD for deep learning workflows deployed on edge devices or across the edge-to-cloud continuum, while accounting for the economic costs, including global network traffic and electricity for edge computing, which depend heavily on the geographical locations of the edge devices. To this end, we proposed a hyperparameter tuning framework, CEDD-optimizer, that leverages key hyperparameters of DD for geographically distributed edge devices to D. Application to Continual Learning maximize the accuracy of downstream deep learning tasks Continual Learning (CL) is a widely used learning paradigm while minimizing total economic cost. We formulated our in which models learn from a continuous stream of datasets. tuning problem in a concrete mathematical form and provided Existing studies [21]–[26] explore the use of DD in CL, which predictive models that are incorporated into our solution is also a potential use case of CEDD-Optimizer. In this use workflow. We thoroughly evaluated our solution across various case, the calibration procedure for the accuracy model can be scenarios and parameter configurations, demonstrating the feasibility of DD in the context of deep learning workflow further optimized by leveraging CL’s nature. deployments on edge computing, where both service quality E. Application to Federated Learning (in terms of deep learning task accuracy) and operating costs Another potential use case of CEDD-optimizer is Federated matter. We hope our study will inspire projects and solutions Learning (FL), while this paper focuses on centralized learning. for the intrinsically coupled challenges between deep learning Several studies employ DD in FL, replacing network parameters performance and economic operating cost. with decentralized updates [13], [27], [28], [70]–[72]. While the ACKNOWLEDGMENTS integration with FL is beyond the scope of this paper, our work also provides the foundation for this extension. Note that the Performance results were partially obtained on systems in concept of price-variation-aware hyperparameter optimization the test environment BEAST (Bavarian Energy Architecture & is valid for DD, regardless of whether the training host is Software Testbed) at the Leibniz Supercomputing Centre. This centralized or distributed. However, the optimization strategy work was supported by the research project Optimierung von needs to be modified for FL due to differences in computation Gasturbinen mit Hilfe von Big Data (AZ-1214-16), funded by and communication patterns across edges. the Bayerische Forschungsstiftung.
R EFERENCES [1] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature, vol. 521, no. 7553, pp. 436–444, 2015. [2] S. AbdulRahman, H. Tout, H. Ould-Slimane, A. Mourad, C. Talhi, and M. Guizani, “A survey on federated learning: The journey from centralized to distributed on-site learning and beyond,” IEEE Internet of Things Journal, vol. 8, no. 7, pp. 5476–5497, 2020. [3] Y. Hu, Z. Yang, C. Zhao, Q. Guo, M. Gao, P. Li, and W. Ji, “Aivd: Adaptive edge-cloud collaboration for accurate and efficient industrial visual detection,” arXiv preprint arXiv:2601.04734, 2026. [4] M. Touhami, M. F. Ahmad Fauzi, Z. Ur Rehman, and S. Mansor, “Federated learning for histopathology image classification: A systematic review,” Diagnostics, vol. 16, no. 1, p. 137, 2026. [5] L. F. R. Moreira, L. N. D. F. Saar, R. Moreira, L. G. F. Rodrigues, B. A. N. Travençolo, and A. R. Backes, “Enabling intelligence on edge through an artificial intelligence as a service architecture,” in 2024 IEEE 13th International Conference on Cloud Networking (CloudNet). IEEE, 2024, pp. 1–8. [6] T. Liu, J. Xie, and H. Dong, “A federated learning with large-small kernel attention network for image classification,” Frontiers in Plant Science, vol. 17, p. 1783587, 2026. [7] Y. Xu, Y. Liao, H. Xu, Z. Wang, L. Wang, J. Liu, and C. Qian, “Fedsnn: Training slimmable neural network with federated learning in edge computing,” IEEE Transactions on Networking, vol. 33, no. 1, pp. 414–429, 2024. [8] S. Elakkiya and K. Thivya, “Comprehensive review on lossy and lossless compression techniques,” Journal of The Institution of Engineers (India): Series B, vol. 103, no. 3, pp. 1003–1012, 2022. [9] Z. Xiong, X. Wu, S. Cheng, and J. Hua, “Lossy-to-lossless compression of medical volumetric data using three-dimensional integer wavelet transforms,” IEEE Transactions on Medical Imaging, vol. 22, no. 3, pp. 459–470, 2003. [10] P. Kavitha, “A survey on lossless and lossy data compression methods,” International Journal of Computer Science & Engineering Technology, vol. 7, no. 03, pp. 110–114, 2016. [11] T. Wang, J.-Y. Zhu, A. Torralba, and A. A. Efros, “Dataset distillation,” arXiv preprint arXiv:1811.10959, 2018. [12] B. Zhao, K. R. Mopuri, and H. Bilen, “Dataset condensation with gradient matching,” arXiv preprint arXiv:2006.05929, 2020. [13] R. Song, D. Liu, D. Z. Chen, A. Festag, C. Trinitis, M. Schulz, and A. Knoll, “Federated learning via decentralized dataset distillation in resource-constrained edge environments,” arXiv preprint arXiv:2208.11311, 2022. [14] K. Wang, B. Zhao, X. Peng, Z. Zhu, S. Yang, S. Wang, G. Huang, H. Bilen, X. Wang, and Y. You, “Cafe: Learning to condense dataset by aligning features,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12 196–12 205. [15] D. Liu, J. Gu, H. Cao, C. Trinitis, and M. Schulz, “Dataset distillation by automatic training trajectories,” arXiv preprint arXiv:2407.14245, 2024. [16] Z. Guo, K. Wang, G. Cazenavette, H. Li, K. Zhang, and Y. You, “Towards lossless dataset distillation via difficulty-aligned trajectory matching,” 2024. [17] Z. Deng and O. Russakovsky, “Remember the past: Distilling datasets into addressable memories for neural networks,” 2022. [18] T. Dong, B. Zhao, and L. Lyu, “Privacy for free: How does dataset condensation help privacy?” in International Conference on Machine Learning. PMLR, 2022, pp. 5378–5396. [19] D. Chen, R. Kerkouche, and M. Fritz, “Private set generation with discriminative information,” Advances in Neural Information Processing Systems, vol. 35, pp. 14 678–14 690, 2022. [20] T. Zheng and B. Li, “Differentially private dataset condensation,” NDSS, 2023. [21] F. Wiewel and B. Yang, “Condensed composite memory continual learning,” in 2021 International Joint Conference on Neural Networks (IJCNN). IEEE, 2021, pp. 1–8. [22] M. Sangermano, A. Carta, A. Cossu, and D. Bacciu, “Sample condensation in online continual learning,” in 2022 International Joint Conference on Neural Networks (IJCNN). IEEE, 2022, pp. 01–08. [23] W. Masarczyk and I. Tautkute, “Reducing catastrophic forgetting with learning on synthetic data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 252– 253.
[24] A. Rosasco, A. Carta, A. Cossu, V. Lomonaco, and D. Bacciu, “Distilled replay: Overcoming forgetting through synthetic samples,” in International Workshop on Continual Semi-Supervised Learning. Springer, 2021, pp. 104–117. [25] E. Yang, L. Shen, Z. Wang, T. Liu, and G. Guo, “An efficient dataset condensation plugin and its application to continual learning,” Advances in Neural Information Processing Systems, vol. 36, pp. 67 625–67 642, 2023. [26] J. Gu, K. Wang, W. Jiang, and Y. You, “Summarizing stream data for memory-restricted online continual learning,” arXiv preprint arXiv:2305.16645, vol. 2, 2023. [27] R. Pi, W. Zhang, Y. Xie, J. Gao, X. Wang, S. Kim, and Q. Chen, “Dynafed: Tackling client data heterogeneity with global dynamics,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 177–12 186. [28] Y. Zhou, X. Ma, D. Wu, and X. Li, “Communication-efficient and attack-resistant federated edge learning with dataset distillation,” IEEE Transactions on Cloud Computing, 2022. [29] B. Zhao and H. Bilen, “Dataset condensation with distribution matching,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 6514–6523. [30] G. Cazenavette, T. Wang, A. Torralba, A. A. Efros, and J.-Y. Zhu, “Dataset distillation by matching training trajectories,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4750–4759. [31] Y. Chen and M. Welling, “Parametric herding,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings, 2010, pp. 97–104. [32] Y. Chen, M. Welling, and A. Smola, “Super-samples from kernel herding,” arXiv preprint arXiv:1203.3472, 2012. [33] M. Toneva, A. Sordoni, R. T. d. Combes, A. Trischler, Y. Bengio, and G. J. Gordon, “An empirical study of example forgetting during deep neural network learning,” arXiv preprint arXiv:1812.05159, 2018. [34] M. Paul, S. Ganguli, and G. K. Dziugaite, “Deep learning on a data diet: Finding important examples early in training,” Advances in Neural Information Processing Systems, vol. 34, pp. 20 596–20 607, 2021. [35] O. Sener and S. Savarese, “Active learning for convolutional neural networks: A core-set approach,” arXiv preprint arXiv:1708.00489, 2017. [36] G. Zhao, G. Li, Y. Qin, and Y. Yu, “Improved distribution matching for dataset condensation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 7856–7865. [37] S. Lee, S. Chun, S. Jung, S. Yun, and S. Yoon, “Dataset condensation with contrastive signals,” in International Conference on Machine Learning. PMLR, 2022, pp. 12 352–12 364. [38] B. Zhao and H. Bilen, “Dataset condensation with differentiable siamese augmentation,” in International Conference on Machine Learning. PMLR, 2021, pp. 12 674–12 685. [39] T. Nguyen, R. Novak, L. Xiao, and J. Lee, “Dataset distillation with infinitely wide convolutional networks,” Advances in Neural Information Processing Systems, vol. 34, pp. 5186–5198, 2021. [40] I. Sucholutsky and M. Schonlau, “Soft-label dataset distillation and text dataset distillation,” in 2021 International Joint Conference on Neural Networks (IJCNN). IEEE, 2021, pp. 1–8. [41] O. Bohdal, Y. Yang, and T. Hospedales, “Flexible dataset distillation: Learn labels instead of images,” arXiv preprint arXiv:2006.08572, 2020. [42] J. Cui, R. Wang, S. Si, and C.-J. Hsieh, “Scaling up dataset distillation to imagenet-1k with constant memory,” 2022. [43] Z. Guo et al., “Towards lossless dataset distillation via difficultyaligned trajectory matching,” in International Conference on Learning Representations (ICLR), 2024. [44] S. Lee et al., “Selmatch: Effectively scaling up dataset distillation via selection-based initialization and partial updates by trajectory matching,” in International Conference on Machine Learning (ICML), 2024. [45] Y. Zhong et al., “Towards stable and storage-efficient dataset distillation: Matching convexified trajectory,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. [46] Y. Chen et al., “Curriculum coarse-to-fine selection for high-ipc dataset distillation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. [47] K. Wang et al., “Distilling datasets into generative models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
[48] J. Su et al., “Dataset distillation via disentangled diffusion model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. [49] T. Renuga et al., “Medsynth: Leveraging generative model for healthcare data sharing,” in International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2024. [50] Anonymous, “Latent video dataset distillation,” OpenReview Preprint, 2025. [51] ——, “On learning representations for tabular data distillation,” arXiv Preprint, 2025. [52] A. Davis, J. Parikh, and W. E. Weihl, “Edgecomputing: extending enterprise applications to the edge of the internet,” in Proceedings of the 13th international World Wide Web conference on Alternate track papers & posters, 2004, pp. 180–187. [53] S. Long, W. Long, Z. Li, K. Li, Y. Xia, and Z. Tang, “A gamebased approach for cost-aware task assignment with qos constraint in collaborative edge and cloud environments,” IEEE Transactions on Parallel and Distributed Systems, vol. 32, no. 7, pp. 1629–1640, 2020. [54] F. Huang, H. Ye, and W. Hao, “Cost-aware resource management based on market pricing mechanisms in edge federation environments,” The Journal of Supercomputing, vol. 79, no. 6, pp. 5939–5961, 2023. [55] H. Li, J. Shen, L. Zheng, Y. Cui, and Z. Mao, “Cost-efficient scheduling algorithms based on beetle antennae search for containerized applications in kubernetes clouds,” The Journal of Supercomputing, vol. 79, no. 9, pp. 10 300–10 334, 2023. [56] S. Rac and M. Brorsson, “Cost-aware service placement and scheduling in the edge-cloud continuum,” ACM Transactions on Architecture and Code Optimization, vol. 21, no. 2, pp. 1–24, 2024. [57] S. Lu, Q. Xia, X. Tang, X. Zhang, Y. Lu, and J. She, “A reliable data compression scheme in sensor-cloud systems based on edge computing,” IEEE Access, vol. 9, pp. 49 007–49 015, 2021. [58] G. Wu, F. Zhou, G. Ding, Q. Wu, and X.-Y. Li, “An efficient heterogeneous edge-cloud learning framework for spectrum data compression,” IEEE Transactions on Mobile Computing, vol. 22, no. 7, pp. 3823–3839, 2023. [59] D. Rosendo, P. Silva, M. Simonin, A. Costan, and G. Antoniu, “E2clab: Exploring the computing continuum through repeatable, replicable and reproducible edge-to-cloud experiments,” in 2020 IEEE International Conference on Cluster Computing (CLUSTER), 2020, pp. 176–186. [60] D. Rosendo, A. Costan, G. Antoniu, M. Simonin, J.-C. Lombardo, A. Joly, and P. Valduriez, “Reproducible performance optimization of complex applications on the edge-to-cloud continuum,” in 2021 IEEE International Conference on Cluster Computing (CLUSTER), 2021, pp. 23–34. [61] D. Rosendo, M. Mattoso, A. Costan, R. Souza, D. Pina, P. Valduriez, and G. Antoniu, “Provlight: Efficient workflow provenance capture on the edge-to-cloud continuum,” in 2023 IEEE International Conference on Cluster Computing (CLUSTER), 2023, pp. 221–233. [62] F. Habeeb, K. Alwasel, A. Noor, D. N. Jha, D. AlQattan, Y. Li, G. S. Aujla, T. Szydlo, and R. Ranjan, “Dynamic bandwidth slicing for time-critical iot data streams in the edge-cloud continuum,” IEEE Transactions on Industrial Informatics, vol. 18, no. 11, pp. 8017–8026, 2022. [63] S. Tang, W. Zhou, L. Chen, L. Lai, J. Xia, and L. Fan, “Batteryconstrained federated edge learning in uav-enabled iot for b5g/6g networks,” Physical Communication, vol. 47, p. 101381, 2021. [64] Y. Mao, J. Zhang, and K. B. Letaief, “Dynamic computation offloading for mobile-edge computing with energy harvesting devices,” IEEE Journal on Selected Areas in Communications, vol. 34, no. 12, pp. 3590–3605, 2016. [65] G. Zhang, W. Zhang, Y. Cao, D. Li, and L. Wang, “Energy-delay tradeoff for dynamic offloading in mobile-edge computing system with energy harvesting devices,” IEEE Transactions on Industrial Informatics, vol. 14, no. 10, pp. 4642–4655, 2018. [66] W. Zhang, S. Li, L. Liu, Z. Jia, Y. Zhang, and D. Raychaudhuri, “Heteroedge: Orchestration of real-time vision applications on heterogeneous edge clouds,” in IEEE INFOCOM 2019 - IEEE Conference on Computer Communications, 2019, pp. 1270–1278. [67] A. Bhattacharjee, A. D. Chhokra, H. Sun, S. Shekhar, A. Gokhale, G. Karsai, and A. Dubey, “Deep-edge: An efficient framework for deep learning model update on heterogeneous edge,” in 2020 IEEE 4th International Conference on Fog and Edge Computing (ICFEC), 2020, pp. 75–84.
[68] H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” 2017. [Online]. Available: https://arxiv.org/abs/1602.05629 [69] J. Wang, Q. Liu, H. Liang, G. Joshi, and H. V. Poor, “Tackling the objective inconsistency problem in heterogeneous federated optimization,” Advances in neural information processing systems, vol. 33, pp. 7611–7623, 2020. [70] Y. Zhou, G. Pu, X. Ma, X. Li, and D. Wu, “Distilled one-shot federated learning,” arXiv preprint arXiv:2009.07999, 2020. [71] J. Goetz and A. Tewari, “Federated learning via synthetic data,” arXiv preprint arXiv:2008.04489, 2020. [72] S. Hu, J. Goetz, K. Malik, H. Zhan, Z. Liu, and Y. Liu, “Fedsynth: Gradient compression via synthetic data in federated learning,” arXiv preprint arXiv:2204.01273, 2022. [73] J. Miano, Compressed image file formats: Jpeg, png, gif, xbm, bmp. Addison-Wesley Professional, 1999. [74] K. Zhao, S. Huang, P. Pan, Y. Li, Y. Zhang, Z. Gu, and Y. Xu, “Distribution adaptive int8 quantization for training cnns,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 4, 2021, pp. 3483–3491. [75] U. Michelucci, “An introduction to autoencoders,” arXiv preprint arXiv:2201.03898, 2022. [76] G. E. Hinton, A. Krizhevsky, and S. D. Wang, “Transforming autoencoders,” in Artificial Neural Networks and Machine Learning–ICANN 2011: 21st International Conference on Artificial Neural Networks, Espoo, Finland, June 14-17, 2011, Proceedings, Part I 21. Springer, 2011, pp. 44–51. [77] L. Deng, “The mnist database of handwritten digit images for machine learning research [best of the web],” IEEE signal processing magazine, vol. 29, no. 6, pp. 141–142, 2012. [78] A. Krizhevsky, G. Hinton et al., “Learning multiple layers of features from tiny images,” 2009. [79] G. Drainakis, K. V. Katsaros, P. Pantazopoulos, V. Sourlas, and A. Amditis, “Federated vs. centralized machine learning under privacyelastic users: A comparative analysis,” in 2020 IEEE 19th International Symposium on Network Computing and Applications (NCA). IEEE, 2020, pp. 1–8. [80] W. P. Review, “Cost of electricity by country 2025,” 2025, accessed: July 8, 2025. [Online]. Available: https://worldpopulationreview.com/ country-rankings/cost-of-electricity-by-country [81] Amazon, “Aws direct connect pricing,” 2025, accessed: July 9, 2025. [Online]. Available: https://aws.amazon.com/directconnect/pricing/ [82] S. Chauhan and A. Rahim, “Introducing aws direct connect sitelink,” 2021. [Online]. Available: https://aws.amazon.com/blogs/networking-and-content-delivery/ introducing-aws-direct-connect-sitelink/ [83] Amazon, “Amazon s3 pricing,” 2025, accessed: July 9, 2025. [Online]. Available: https://aws.amazon.com/s3/pricing/ [84] T. Domhan, J. T. Springenberg, F. Hutter et al., “Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves.” in IJCAI, vol. 15, 2015, pp. 3460–8. [85] A. Klein, S. Falkner, J. T. Springenberg, and F. Hutter, “Learning curve prediction with bayesian neural networks,” in International conference on learning representations, 2017. [86] M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Poczos, R. R. Salakhutdinov, and A. J. Smola, “Deep sets,” Advances in neural information processing systems, vol. 30, 2017. [87] K. Krishnamoorthy, Handbook of Statistical Distributions with Applications, 2nd ed. Boca Raton, FL: Chapman and Hall/CRC, 2016. [88] SciPy, “Optimization and root finding (scipy.optimize),” 2008, accessed: July 21, 2025. [Online]. Available: https://docs.scipy.org/doc/scipy/ reference/optimize.html [89] ——, “scipy.optimize.linear sum assignment,” 2026, accessed: Jan 13, 2026. [Online]. Available: https://docs.scipy.org/doc/scipy/reference/ generated/scipy.optimize.linear sum assignment.html [90] D. F. Crouse, “On implementing 2d rectangular assignment algorithms,” IEEE Transactions on Aerospace and Electronic Systems, vol. 52, no. 4, pp. 1679–1696, 2016. [91] TQ-Systems, “Tqmx80uc x86 family,” 2020, accessed: July 21, 2025. [Online]. Available: https://www.tq-group.com/filedownloads/files/products/embedded/ data-sheets/x86/embedded-modul/COM-Express-Compact/ TQMx80UC/EMB DB TQMx80UC EN Rev109 Web 01.pdf
[92] H. Xiao, K. Rasul, and R. Vollgraf, “Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms,” arXiv preprint arXiv:1708.07747, 2017. [93] Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng et al., “Reading digits in natural images with unsupervised feature learning,” in NIPS workshop on deep learning and unsupervised feature learning, vol. 2011, no. 2. Granada, 2011, p. 4. [94] F. Fastai/imagenette, “A smaller subset of 10 easily clas- sified classes from imagenet, and a little more french.” https://github.com/fastai/ imagenette, 2013. [95] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009, pp. 248–255. [96] S. Hong, S. Jang, W. Kweon, S. Kim, G. Lee, and H. Yu, “Harmonic dataset distillation for time series forecasting,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 26, 2026, pp. 21 770–21 778. [97] S. Gidaris and N. Komodakis, “Dynamic few-shot visual learning without forgetting,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4367–4375. [98] D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Instance normalization: The missing ingredient for fast stylization,” arXiv preprint arXiv:1607.08022, 2016. [99] J. Du, Y. Jiang, V. T. Tan, J. T. Zhou, and H. Li, “Minimizing the accumulated trajectory error to improve dataset distillation,” arXiv preprint arXiv:2211.11004, 2023. [100] B. Zhao and H. Bilen, “Synthesizing informative training samples with gan,” arXiv preprint arXiv:2204.07513, 2022. [101] G. Li, R. Togo, T. Ogawa, and M. Haseyama, “Dataset distillation using parameter pruning,” arXiv preprint arXiv:2209.14609, 2022. [102] Y. Liu, Z. Li, M. Backes, Y. Shen, and Y. Zhang, “Backdoor attacks against dataset distillation,” arXiv preprint arXiv:2301.01197, 2023. [103] Z. Wang, Y. Xu, C. Lu, and Y.-L. Li, “Dancing with still images: Video distillation via static-dynamic disentanglement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6296–6304. [104] S. Edelkamp and S. Schrödl, Heuristic search: theory and applications. Elsevier, 2011. [105] L. Bottou, “Stochastic gradient descent tricks,” in Neural networks: tricks of the trade: second edition. Springer, 2012, pp. 421–436.
[106] S.-i. Amari, “Backpropagation and stochastic gradient descent method,” Neurocomputing, vol. 5, no. 4-5, pp. 185–196, 1993. [107] R. H. Byrd, R. B. Schnabel, and G. A. Shultz, “A trust region algorithm for nonlinearly constrained optimization,” SIAM Journal on Numerical Analysis, vol. 24, no. 5, pp. 1152–1170, 1987. [108] SciPy Community, “scipy.optimize.minimize with the trust-constr method,” SciPy v1.18.0 documentation, 2026. [Online]. Available: https://docs.scipy.org/doc/scipy/reference/optimize. minimize-trustconstr.html [109] R. Storn and K. Price, “Differential evolution—a simple and efficient heuristic for global optimization over continuous spaces,” Journal of Global Optimization, vol. 11, pp. 341–359, 1997. [110] SciPy Community, SciPy v1.18.0 documentation, 2026. [Online]. Available: https://docs.scipy.org/doc/scipy/reference/generated/scipy. optimize.differential evolution.html [111] L. Beal, D. Hill, R. Martin, and J. Hedengren, “Gekko optimization suite,” Processes, vol. 6, no. 8, p. 106, 2018. [112] P. Belotti, C. Kirches, S. Leyffer, J. Linderoth, J. Luedtke, and A. Mahajan, “Mixed-integer nonlinear optimization,” Acta Numerica, vol. 22, pp. 1–131, 2013. [113] N. Carlini, V. Feldman, and M. Nasr, “No free lunch in” privacy for free: How does dataset condensation help privacy”,” arXiv preprint arXiv:2209.14987, 2022. [114] X. Wang, S. Sha, and Y. Sun, “Dcfl: Non-iid awareness dataset condensation aided federated learning,” in 2024 International Joint Conference on Neural Networks (IJCNN). IEEE, 2024, pp. 1–8. [115] C.-Y. Huang, R. Jin, C. Zhao, D. Xu, and X. Li, “Federated virtual learning on heterogeneous data with local-global distillation,” arXiv preprint arXiv:2303.02278, 2023. [116] Y. Wang, H. Fu, R. Kanagavelu, Q. Wei, Y. Liu, and R. S. M. Goh, “An aggregation-free federated learning for tackling data heterogeneity,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 233–26 242. [117] Y. Jia, S. Vahidian, J. Sun, J. Zhang, V. Kungurtsev, N. Z. Gong, and Y. Chen, “Unlocking the potential of federated learning: The symphony of dataset distillation via deep generative latents,” arXiv preprint arXiv:2312.01537, vol. 2, 2023.