Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling Jonathan Bader∗ , Edgar Blumenthal∗ , Marten Eckardt∗ , Justus Krebs∗ , Joel Witzke∗ , Xemena Wysokinska∗ , Haci Ismail Aslan, and Odej Kao
arXiv:2604.18043v1 [cs.DC] 20 Apr 2026
TU Berlin, Germany {jonathan.bader, joel.witzke, aslan, odej.kao}@tu-berlin.de {e.blumenthal, m.eckardt, justus.krebs, x.wysokinska}@campus.tu-berlin.de
Abstract—In modern distributed systems, efficient resource allocation is a vital aspect to maintain scalability, reduce operational costs, and ensure fast execution even across heterogeneous workloads. Predictive models for resource usage are essential tools for optimizing allocation and preventing system bottlenecks. Predictive memory allocation has asymmetric costs as a key challenge: underallocation causes failures while overallocation wastes memory. We propose a regression method based on a LightGBM and XGBoost ensemble trained to predict high conditional quantiles. To further account for the high cost of underallocations we add a multiplicative safety factor. With our method we are able to reduce the number of under-allocated jobs from 4.17% to 2.89% and average overallocation from 148% to 44.51% on a real-world dataset of build jobs provided by SAP. We further explore the pareto frontier between optimization for underallocation and for overallocation. Index Terms—memory prediction, resource prediction, machine learning, resource management
I. I NTRODUCTION Distributed clusters are used by many companies and institutions on a daily basis. Because such architectures are the backbone of modern large-scale computing systems, correct resource allocation is a high priority. One such example is memory allocation. This is challenging: underallocation, assigning too little memory, often causes crashes, whereas excessive overallocation wastes resources. The task of memory allocation has evolved from custom and general-purpose allocators [1], through dynamic schemes (e.g. binning [2]) to machine learning-based optimization [3]. Other related approaches involve recurrent neural networks [4], LSTM based models [5] and dynamic ensemble learning for scientific workflows [6]. For this study, we use a dataset from SAP containing detailed traces of build jobs1 . Analyzing baseline memory allocation done through manual memory limitations by SAP developers revealed opportunities to reduce both under- and overallocation, motivating the search for and development of a better approach. 1 https://github.com/SAP/task-execution-data-set
*equal contribution, alphabetical order
We evaluated multiple classification and regression approaches. The best results were achieved using a quantile regression ensemble [7] combining XGBoost [8] and LightGBM [9]. This approach achieved more than a two-thirds reduction in memory waste and reduced underallocation to below 3%, representing a 50% improvement over the SAP baseline. II. E VALUATION A. Dataset Features The dataset contains Continuous Integration (CI) infrastructure metadata on a cluster. The data includes resource usage information for build jobs. After anonymization, each build has 19 feature columns with the most relevant being: • time: timestamp of the build. • memory fail count: kernel memory allocation failure count during the build; used to identify whether a job was under-allocated. • buildProfile: contains informaton about the architecture, compiler, and the optimization level. • makeType: build type indicator. • max rss: peak memory usage during the build (bytes). • memreq: requested memory for the build container (MB). We derived an additional 21 predictive features from the initial 19 raw telemetry columns through feature engineering (excluding one-hot encodings). We applied temporal decomposition to capture daily and weekly seasonality in resource usage [10], build profile parsing, workload characterization and lagged statistical history. After creating our model, the five most important features were lag 1 grouped (memory usage of the direct predecessor job), rolling p95 rss g1 w5 (rolling 95th percentile of memory usage for similar jobs), jobs, branch id str and ts weekofyear. This dominance of historical features validates our feature engineering strategy and demonstrates that a job’s recent, localized memory history is the most critical predictor for its future memory needs. B. Hyperparameter Search We used Bayesian optimization to tune each model’s hyperparameters using the Optuna library [11]. Bayesian optimization builds a surrogate model of performance (e.g., a Gaussian process) and selects new hyperparameters via an acquisition
©2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. DOI: https://doi.org/10.1109/BigData66926.2025.11402160
total predicted rss(θ) θ total actual max rss where θ are the parameters of the respective model. We penalized underallocation errors since they could lead to job failures, wasting previous computation and necessitating a rerun. The penalty factor was set to 5 based on empirical tuning and represents our trade-off between prioritizing job success while maintaining overall resource efficiency. Each trial consisted of 3-fold cross-validation to ensure that the chosen hyperparameters generalize across folds and do not overfit a single train/validation split. We pooled all outof-fold (OOF) predictions and their corresponding true values from each fold to calculate Cost(θ) once on the aggregated predictions. This gives equal weight to every test instance, providing a more robust performance estimate than averaging the costs of each fold [14], since the distribution of target values and the size of each fold can vary significantly. min Cost(θ) = 5 · under alloc(θ) +
C. Regression Ensemble Setup To further reduce underallocation, we predicted the upper α-quantile (with α ∈ [0.90, 0.99]) of memory instead of the mean. This approach trains the model to produce predictions that are expected to be higher than the true value in α% of cases, naturally reducing the risk of underallocation [15]. We evaluated both single quantile regressors and multiple two-model quantile ensembles. Ensembles of models can reduce correlation between learners and often generalize more robustly on tabular data [16]. Each of the base learners a and b predicts an α-quantile, and the ensemble output is the per-row maximum, scaled by a safety factor s ∈ [1.00, 1.15]: ŷ(x) = max ŷα(a) (x), ŷα(b) (x) · s The per-row maximum aggregation method was selected for its simplicity and effectiveness in enforcing a conservative allocation strategy. This approach ensures that for any given job, the allocation is guided by the more cautious of the two models, thereby minimizing the risk of underallocation. While we explored other techniques, such as mean and percentilebased aggregation [17], as well as expanding the ensemble with more models, these alternatives failed to deliver better performance and introduced unnecessary complexity. Therefore, maximum aggregation with two quantile loss models provided the optimal balance of caution, accuracy, and simplicity.
Pareto Frontier (Waste vs Underallocation) Pareto frontier Low Waste ( =0.90, s=1.00) Low Underallocation ( =0.99, s=1.15) Balanced ( =0.94, s=1.11)
12
Underallocation (%)
policy (e.g., expected improvement). Here, the TPE sampler was used [12] in accordance to Snoek et al. [13], balancing high predicted performance (exploitation) against uncertainty (exploration) and makes use of past evaluations for guidance, often finding good hyperparameter settings in fewer trials than simple grid or random search. The search space included generic parameters (e.g. learning rate, number of trees/estimators) and model-specific parameters such as the quantile parameter for the ensembles. The cost function used for training the models balances the amount of under-allocated jobs (%) against the total overallocation ratio to ensure predictions are sufficient but not wasteful:
10 8 6 4 2 0 30
40
50
60
70
Waste (Total Over-allocation %)
Fig. 1. Pareto frontier for our LightGBM+XGBoost ensemble model with points corresponding to different (α, s) pairs.
In our regression experiments we focused on gradientboosted tree models: Scikit-learn’s GradientBoosting (GB), XGBoost (XGB), LightGBM (LGB), and CatBoost (Cat). We evaluated the following model families: • Heterogeneous Ensembles: Various pairwise combinations of the four base learners (e.g., LGB+XGB, GB+LGB). • Homogeneous Ensembles: Combinations of two identical learners, each with an independent hyperparameter search (e.g., XGB+XGB, Cat+Cat). • Single-Model Baselines: Standalone LightGBM and XGBoost models using their native quantile objectives. Heterogeneous pairs (e.g., LGB+XGB) provide diversity across libraries, which reduces error correlation without leaving the quantile framework. This focus also aligns with benchmarks showing that gradient-boosted models are consistently state-of-the-art for structured tabular data [18], [19]. D. Exploring the Pareto Frontier and Results By adjusting the model’s quantile level (α) and safety multiplier (s), we can create different versions of the model to prioritize either safety or efficiency. This creates a pareto frontier between optimizing for under- vs. for overallocation shown in Fig. 1. Having different models on that frontier allows for flexibility and can serve different operational needs or business priorities: • Balanced: Our best performing model based on our defined cost function, offering a strong balance between safety and efficiency (2.89% underallocation, 44.51% overallocation). • Low Waste: An aggressive model that minimizes overallocation (25.5%) at the cost of a higher underallocation rate (12.6%). • Low Under-allocation: A very conservative model that nearly eliminates underallocation (0.20%) at the cost of
Memory Allocation Comparison
100
3.8% 3.9%
22.5%
12.7%
80
Share of Jobs (%)
14.7%
R EFERENCES
60 30.6% 76.6%
40
20
28.0%
0 Allocation Category
4.2% Baseline
Under-allocated Well-allocated (1x-2x) Severely Over (2x-3x) Extremely Over (3x-4x) Massively Over (4x+)
Further funding received by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) as FONDA (Project 414984028, SFB 1404). We also thank SAP for their data and friendly cooperation.
Allocation Strategy
2.9% lgb_xgb_ensemble
Fig. 2. Job distribution by allocation quality for our method vs. the baseline.
higher resource waste (78.2%). That is still substantially lower than the baselines 148% on the same hold-out set. Fig. 2 shows the performance in different levels of overallocation of our balanced approach compared to the baseline. With inference taking about 2 · 10−5 seconds it is also suitable for predictions in realtime scenarios. III. C ONCLUSION In this paper we addressed the challenge of memory allocation dealing with asymmetric cost as under-provisioning is significantly more costly than over-provisioning. Our proposed quantile regression ensemble reduced underallocation from the baseline’s 4.17% to 2.89% and, critically, decreased total overallocation (resource waste) by over 70% from 148% to 44%. A key finding from our experiments is that a conservative architectural design was more influential than the specific choice of algorithm. The consistent high performance across different regression ensemble combinations indicates that the core strategy of using a high-quantile objective and max aggregation was the primary driver of success. Future work could extend this research, e.g., by exploring a hybrid policy that uses an additional conservative model for build jobs predicted to require large amounts of memory, where underallocation implications are the highest. Another approach could use reinforcement learning to learn an optimal allocation online, dynamically adapting to changing workloads and current cluster utilizations in real-time. Our code, including an application for realtime predictions and trained models, can be found on GitHub2 . ACKNOWLEDGMENTS This paper was partially supported by the Swarmchestrate project of the European Union’s Horizon 2023 Research and Innovation programme under grant agreement no. 101135012. 2 https://github.com/Zmart64/SAPResourceOptimizer
[1] E. D. Berger, B. G. Zorn, and K. S. McKinley, “Reconsidering custom memory allocation,” in Proceedings of the 17th ACM SIGPLAN Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA), 2002, pp. 1–12. [2] A. Wolke, B. Tsend-Ayush, C. Pfeiffer, and M. Bichler, “More than bin packing: Dynamic resource allocation strategies in cloud data centers,” Information Systems, vol. 52, pp. 83–95, 2015. [3] C. Witt, M. Bux, W. Gusew, and U. Leser, “Predictive performance modeling for distributed batch processing using black box monitoring and machine learning,” Information Systems, vol. 82, pp. 33–52, 2019. [4] W. Zhang, B. Li, D. Zhao, F. Gong, and Q. Lu, “Workload prediction for cloud cluster using a recurrent neural network,” in 2016 International Conference on Identification, Information and Knowledge in the Internet of Things (IIKI), 2016, pp. 104–109. [5] K. Thonglek, K. Ichikawa, K. Takahashi, H. Iida, and C. Nakasan, “Improving resource utilization in data centers using an lstm-based prediction model,” in 2019 IEEE International Conference on Cluster Computing (CLUSTER), 2019, pp. 1–8. [6] J. Bader, F. Skalski, F. Lehmann, D. Scheinert, J. Will, L. Thamsen, and O. Kao, “Sizey: Memory-efficient execution of scientific workflow tasks,” in 2024 IEEE International Conference on Cluster Computing (CLUSTER), 2024, pp. 370–381. [7] N. Meinshausen, “Quantile regression forests,” Journal of Machine Learning Resesearch, vol. 7, pp. 983–999, 2006. [8] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 785–794. [9] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Liu, “LightGBM: A highly efficient gradient boosting decision tree,” in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 3146–3154. [10] K. A. Huang, W. M. Hardin, N. S. Prakash, and W. Hardin, “Forecasting emergency room patient volumes using extreme gradient boosting with temporal and seasonal feature engineering: A comparative study across hospitals,” Cureus, vol. 17, no. 6, 2025. [11] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A nextgeneration hyperparameter optimization framework,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019, pp. 2623–2631. [12] S. Watanabe, “Tree-structured parzen estimator: Understanding its algorithm components and their roles for better empirical performance,” 2025. [Online]. Available: https://arxiv.org/abs/2304.11127 [13] J. Snoek, H. Larochelle, and R. P. Adams, “Practical bayesian optimization of machine learning algorithms,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 25, 2012. [14] G. Forman and M. Scholz, “Apples-to-apples in cross-validation studies: pitfalls in classifier performance measurement,” SIGKDD Explorations Newsletter, vol. 12, no. 1, pp. 49–57, Nov. 2010. [15] R. Koenker and G. Bassett, “Regression quantiles,” Econometrica, vol. 46, no. 1, pp. 33–50, 1978. [16] R. Caruana, A. Niculescu-Mizil, G. Crew, and A. Ksikes, “Ensemble selection from libraries of models,” in Proceedings of the Twenty-First International Conference on Machine Learning (ICML), 2004, p. 18. [17] F. Busetti, “Quantile aggregation of density forecasts,” Oxford Bulletin of Economics and Statistics, vol. 79, no. 4, pp. 495–512, 2017. [18] L. Grinsztajn, E. Oyallon, and G. Varoquaux, “Why do tree-based models still outperform deep learning on typical tabular data?” in Advances in Neural Information Processing Systems (NeurIPS), vol. 35. Curran Associates, Inc., 2022, pp. 507–520. [19] R. Shwartz-Ziv and A. Armon, “Tabular data: Deep learning is not all you need,” Information Fusion, vol. 81, pp. 84–90, 2022.