Distributed Learning as a Service: The Developer’s Perspective Tianyue Chu1 , Filippo Vannella1 , Dimitra Tsigkari1 , Paula Delgado-Santos1 Fernando López1 , Pablo Gomez Guerrero1 , Sotirios Spantideas2 , David Solans Noguero1 1
arXiv:2609.31061v1 [cs.LG] 25 Sep 2026
2
Telefónica Scientific Research, Telefónica, Madrid, Spain National and Kapodistrian University of Athens, Psachna, Evia, Greece
Abstract—Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed Learning as a Service) from the developer’s vantage point. Using a single admin dashboard, the developer initiates a distributed/federated learning job and is able to activate Differential Privacy (DP), Split Learning (SL), Hierarchical Aggregation (HA), and Knowledge Distillation (KD) as declarative options, with no change to the clients’ code. We demonstrate the complete service lifecycle on an industrial smart-home Wake-up Word (WuW) task, using the “Ok Aura” dataset. Once the developer initiates a distributed learning job by toggling DP, SL, HA, and KD in the admin dashboard, the system dispatches the job to a set of Android clients and Dockerized helper aggregators. In the demonstration, these mechanisms run live across configurations. Then, the clients train the model locally and return their updates. The trained model is served to a consumer-side Android application that performs on-device WuW detection on a live microphone stream. In particular, the conference attendees will be invited to speak the trigger phrase and monitor in real time the per-class confidence and inference latency. Finally, we release the source code and short video walkthroughs of these configurations. Index Terms—Distributed learning, federated learning, edge computing, MLOps
I. Introduction Machine Learning (ML) is increasingly trained on Internet of Things (IoT) and other edge devices, close to where data is produced. Such data often cannot be collected in one place, because it is private or costly to move. A common response is distributed learning, which spreads the training across many devices rather than one server. Federated Learning (FL) is a widely used method of distributed learning: each device keeps its own data and shares only model updates [1]. Federated Learning as a Service (FLaaS) [2], [3] is a service that allows an application developer to manage the FL training without orchestrating the underlying system. Deploying FL as a production service, however, raises four challenges a basic FL loop does not address. First, the exchanged updates may still reveal private information. Second, devices with limited memory may be unable to train large models. Third, a single aggregator limits scalability as more clients join. Fourth, transmitting a large model in every round is costly
2 Central Server
Notification Path
Notification Service
4
3 1
Distributed resource pool
Device Workers
Job Core Logic & Model Management spec Central DP Admin interface Split Learning Knowledge Distillation Dashboard Model Aggregation User/workflow Device Scheduling management Load balancing
Admin & Front-End
Helper node A Helper node B
Helper node H
Helpers
Push Notification System Local Training Local DP
Local Inference Apps
Database
Server Cluster
Client Device
Fig. 1. DLaaS architecture annotated for the demonstration as a job flows through the system: (1) the developer composes the job at the admin frontend (DP mechanism and its (ε, δ) values, the SL, HA, and KD toggles) and submits the specifications; (2) the server cluster compiles the given specs into a per-round plan and runs the server-side core logic (central DP, SL, KD, model aggregation, and device scheduling), driving the distributed learning round loop; (3) the helper tier performs partial aggregation over client clusters and can host the SL server-side part; and (4) each client device runs local training, local DP, and sends the trained model back to the server.
over constrained links. In a real deployment, such as a smart home, these challenges arise together. DLaaS (Distributed Learning as a Service) overcomes these obstacles by extending the open-source FLaaS framework [2], [3] with four mechanisms: Differential Privacy (DP), Split Learning (SL), Hierarchical Aggregation (HA), and Knowledge Distillation (KD). Each of these mechanisms appears as a declarative switch in the same job specification rather than a separate system to integrate. To the best of our knowledge, DLaaS is the first framework to jointly support all four mechanisms while preserving access to the FL functionalities of FLaaS. This demo presents DLaaS from the developer’s perspective, showing what it is like to configure, launch, and watch a distributed learning job running from end to end. We demonstrate DLaaS on an industrial smart-home Wakeup Word (WuW) detection workload, the “Ok Aura” task [4], in which an always-on device must recognize a trigger phrase on-device. The WuW detection task is a representative edgelearning workload since it combines privacy-sensitive audio data, continuously operating resource-constrained devices, and large-scale deployment requirements. The demonstration shows how privacy, device-awareness, aggregation scale, and communication cost can be turned on and off as first-class service policies. To make the walkthrough reproducible beyond
© IFIP, 2026. This is the author’s version of a work accepted for publication in the 22nd International Conference on Network and Service Management (CNSM 2026).
(a) create and manage projects
(b) select model, dataset, and FL rounds
(c) configure privacy and SL / HA / KD modules
Fig. 2. (1) Compose — the server-side admin interface (Django admin). (a) the admin home, where projects are created, scheduled, and managed; (b) a project’s training configuration, selecting the model, dataset, and number of FL rounds; and (c) the DP mechanism with its (ε, δ) configuration and the SL, HA, and KD module toggles. The whole job is composed from this one interface.
the demo, we also publish the source code and a short video extract for each different configuration. II. DLaaS Overview Fig. 1 presents the DLaaS architecture. It has three tiers: (i) a central coordinator (admin front-end (1) and server cluster (2) in the figure), (ii) a helper tier (3), and (iii) an on-device client SDK (4). The coordinator, a Django REST service, exposes the control plane and holds the federated state. The helper tier is a set of stateless Dockerized FastAPI aggregators that perform partial aggregation. The client SDK runs local training and inference on Android. Through the control plane, a job is declared rather than coded. The developer selects the dataset and training setup, a DP mechanism with its (ε, δ) budget, FL or SL execution, a flat or hierarchical topology, and toggles distillation. Given these choices, the coordinator compiles the specification into a per-round plan and pushes it to the helpers and clients. Depending on which modules are active, the data plane then carries the model weights, the intermediate activations (under SL), or the distilled updates. We now elaborate on the four modules of DLaaS. (i) DP: DLaaS supports central and local DP. The central DP clips per-client updates and adds calibrated Gaussian noise at the server, with the noise multiplier derived offline from the target (ε, δ) under Rényi-DP composition [5] using the Opacus subsampled-Gaussian accountant [6]. Local DP privatizes each client’s update on-device via the analytic Gaussian mechanism [7], removing the trusted-aggregator assumption at the cost of higher noise. (ii) SL: for devices with limited memory, DLaaS splits the model at a configurable cut layer, so the client computes the frozen first part and streams activations to a server (or helper) that trains the remaining layers [8]. (iii) HA: a helper tier performs weighted partial aggregation over client clusters and forwards a single result upstream, reducing server-side fan-in from O(N ) to O(H) while preserving FedAvg equivalence [9]; helpers can additionally host the SL server-side part, so SL and HA compose. (iv) KD: DLaaS distills a high-capacity teacher into a compact student server-side and ships only the student to clients [10], reducing per-round model transmission and on-device footprint while leaving the client training loop unchanged. All four are exposed as independent fields of the job specification and can be enabled together in one job.
$ python manage.py tick (project 87, round 0) (a) Reported models are enough (Ratio 1.00). Round ’0’ is complete. Creating next round: 1 (b) Using helper. Submitting group 1/1 to helper on port 8500 [NET] uplink_bytes=8000259 downlink_bytes=8000189 latency_s=0.579 [DEBUG] Helper container logs: INFO: Uvicorn running on http://0.0.0.0:8500 (c) INFO: "POST /aggregate HTTP/1.1" 200 OK Helper 1 (port 8500, size 1) finished in 3.72 s Helper-based aggregation complete. Model weights saved to projects/87/1/model_weights.bin (d) Sending train request: users:[’test_user1’] data: {’type’: ’train’, ’project’: 87, ’round’: 1, ’trainingMode’: ’BASELINE’, ’localDP’: 0, ...} Push sent successfully.
Fig. 3. (2) Dispatch & aggregate — excerpt of the system coordinator terminal, with the relevant messages marked: (a) round 0 is complete; (b) the round is dispatched to a helper container; (c) the helper aggregates and the model is saved; (d) round 1 is pushed to the clients.
III. Demonstration The demonstration walks developers through the full DLaaS service loop on the WuW workload, with the live screen for each step shown in Figs. 2–5. A job is composed once (step 1); the system coordinator and clients then repeat steps 2 and 3 for T federated rounds before the consumer app downloads the trained global model and runs live “Ok Aura” detection (step 4). In the demo, one complete FL round (local training followed by global aggregation) runs live. Reaching a converged model takes many such rounds, so the model served to the consumer app in step 4 is taken from a full training run performed beforehand. (1) Compose. From the server-side admin interface, the developer schedules a project, selecting the model, dataset, and number of FL rounds, and configures privacy and modules: the DP mechanism (central or local) and budget ε, and the SL, HA (with H helpers), and KD toggles (Fig. 2). (2) Dispatch and aggregate. The coordinator compiles the spec into a fixed per-round plan and drives the round loop: it recruits clients, dispatches the round, and the helper tier aggregates the returned updates into the new global model before pushing the next round (Fig. 3). Toggling HA on the same project makes its effect directly observable at the coordinator: the server then receives one aggregated result per helper rather than every client update, so the per-round fan-in drops from the client count to the number of helpers H. (3) Local training. Each client runs the FL round on an Android device (a Pixel profile on an arm64-v8a emulator), training the model head on its local samples and returning the
Fig. 4. (3) Local training — the Android client. Left: the authenticated app on the emulator. Right: the Logcat lines reporting per-epoch training loss and accuracy (timestamp, PID, and tag columns omitted for readability).
update; the Logcat reports the per-epoch loss and accuracy (Fig. 4). Under SL it instead exports bottleneck activations, and under local DP it adds noise to the update on-device. (4) Detect. After the rounds complete, the trained global model is served to a consumer-side Android app that runs continuous on-device detection; an attendee can say “Ok Aura” and see the per-class confidence and an inference latency of only a few milliseconds (Fig. 5), closing the loop from configuration to a working consumer experience. Evaluation of DLaaS. Since a few live rounds cannot fully showcase the performance of DLaaS, the quantitative tradeoffs below come from our offline evaluation on the same testbed, on CIFAR-10 and on the WuW task. There, we observe that SL trims peak on-device memory by 33%; KD shrinks the broadcast model by 88.8% on CIFAR-10 and 92.4% on the WuW task, a 78.5 MB teacher distilled to a 5.9 MB student; and HA replaces the server’s O(N ) perround fan-in with O(H), one helper projected to serve roughly 139 clients within an 8 GB budget. To surface these effects, especially the SL and KD savings, the demonstration runs the heavier CIFAR-10 model.
Fig. 5. (4) Detect — the consumer app. It downloads the trained global model and reports live on the “Ok Aura” detection task with per-class confidence and inference latency.
configurable operational primitives, DLaaS paves the way towards production-grade distributed learning systems. Acknowledgments This research is supported by the European Union’s Horizon Europe research and innovation actions under grant agreement No 101168560 (CoEvolution), the European Union under TaRDIS (GA 101093006), and the Horizon MSCA Postdoctoral Fellowship OPALS (grant agreement 101210495). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them.
IV. Demo Setup and Usage The demonstration runs self-contained on a single laptop that hosts the coordinator, the Dockerized helper aggregators, and the emulated Android client, together with one Android phone running the WuW inference application. Network connectivity is required, since the coordinator sends round notifications to the client devices. The code is open source,1 and we provide videos of a complete run for those configurations: • FL training: https://youtu.be/zvImEQucwX8 • DP: https://youtu.be/H4SA9UqQjOM • SL: https://youtu.be/6CqTziGy3J4 • HA: https://youtu.be/XeBnISOqaSw • KD: https://youtu.be/XiZS16lx-2s • WuW training: https://youtu.be/ppHJI2smsC8 • WuW inference: https://youtu.be/QqXN84qxETo V. Conclusion This demonstration puts a developer in control of DLaaS and shows that four capabilities usually added through separate tooling, privacy, device-awareness, scalability, and communication efficiency, can instead be exposed as composable service-level policies within a distributed learning platform. By transforming advanced distributed learning techniques into 1 https://github.com/Telefonica-Scientific-Research/DLaaS-Server
References [1] H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Proc. of AISTATS, 2017. [2] N. Kourtellis, K. Katevas, and D. Perino, “FLaaS: Federated learning as a service,” in Proc. 1st Workshop on Distributed Machine Learning (DistributedML), 2020, arXiv:2011.09359. [3] K. Katevas, D. Perino, and N. Kourtellis, “FLaaS – enabling practical federated learning on mobile environments,” in Proc. 20th Annual Int. Conf. on Mobile Systems, Applications and Services (MobiSys, Demo). ACM, 2022, pp. 605–606. [4] F. López Gavilánez, J. Luque Serrano, and G. G. Pablo, “Okey aura wake-up word dataset (test wuwdc 2024),” Aug. 2024. [Online]. Available: https://doi.org/10.5281/zenodo.13601867 [5] I. Mironov, “Rényi differential privacy,” in 2017 IEEE 30th computer security foundations symposium (CSF). IEEE, 2017, pp. 263–275. [6] A. Yousefpour, I. Shilov, A. Sablayrolles, D. Testuggine, K. Prasad, M. Malek, J. Nguyen, S. Ghosh, A. Bharadwaj, J. Zhao, G. Cormode, and I. Mironov, “Opacus: User-friendly differential privacy library in PyTorch,” arXiv preprint arXiv:2109.12298, 2021. [7] C. Dwork and A. Roth, “The algorithmic foundations of differential privacy,” Foundations and trends® in theoretical computer science, vol. 9, no. 3-4, pp. 211–487, 2014. [8] C. Thapa, P. C. M. Arachchige, S. Camtepe, and L. Sun, “Splitfed: When federated learning meets split learning,” in Proceedings of AAAI, vol. 36, no. 8, 2022, pp. 8485–8493. [9] L. Liu, J. Zhang, S. H. Song, and K. B. Letaief, “Client-edge-cloud hierarchical federated learning,” in Proc. of IEEE ICC, 2020, pp. 1–6. [10] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NeurIPS Deep Learning and Representation Learning Workshop, 2015.