Peak-Aware Short-Term Load Forecasting Across Distribution Grid Aggregation Levels Souhardya Chattopadhyay1,2† , Julian Oelhaf1†* , Antonia Schoening2 , Jessica Deuschel2 , Bitan Bhattacharyya2 , Christian Bergler3 , Andreas Maier1 , Siming Bayer1 1
arXiv:2609.18588v1 [cs.LG] 16 Sep 2026
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg 2 Siemens AG, Smart Infrastructure 3 Department of Electrical Engineering, Media and Computer Science, Ostbayerische Technische Hochschule Amberg-Weiden Erlangen, Germany * Corresponding author: [email protected] Abstract—For distribution system operators, short-term load forecasting (STLF) supports congestion management, voltage control, and asset protection. Most existing approaches focus on overall accuracy across all time steps and neglect performance during high-demand (HD) periods, where larger forecast errors can increase the risk of congestion and voltage violations. In this paper, we study peak-aware STLF across three operator-relevant distribution grid aggregation levels, area codes (AC), secondary substations (SUB), and low-voltage (LV) feeders, using open datasets from the United Kingdom and Switzerland. We compare statistical baselines, machine learning models (LightGBM and XGBoost), and recent time-series foundation models (ChronosBolt and Chronos-2) under a peak-aware evaluation framework that reports both overall and HD forecasting performance using NMAE and MAPE. The results show that Chronos-2 achieves the best HD performance across all aggregation levels, with HDNMAE and HD-MAPE of 0.039 and 4.53 % at AC, 0.080 and 9.45 % at SUB, and 0.138 and 16.14 % at LV, while Chronos-Bolt consistently ranks second best. Compared with the gradientboosted ML models, Chronos-2 reduces mean HD-NMAE by about 20–51 % across levels while remaining best or near-best on the overall metrics. A quantile analysis of the probabilistic Chronos outputs further identifies aggregation-specific operating points, and runtime measurements indicate that foundationmodel inference is fast enough for practical deployment. Overall, the findings highlight peak-aware evaluation and aggregationspecific quantile selection as a practical pathway toward more operationally relevant STLF in distribution networks. Index Terms—Short-Term Load Forecasting, Peak-Aware Forecasting, Distribution Grid Networks, Aggregation Levels, Time-Series Foundation Models
I. I NTRODUCTION Energy systems are becoming harder to operate due to transport electrification, electric heating, distributed renewable generation, and rising demand variability at the grid edge. While these trends support decarbonization, they also increase short-term local fluctuations as electrified end uses such as electric vehicles (EVs), heat pumps become more widespread and create sharper coincident peaks [1]–[3]. For distribution system operators (DSOs), this is especially critical at lower grid levels, where short demand surges can trigger congestion, †These authors contributed equally to this work.
voltage deviations, and asset stress [2], [3]. Hence, shortterm load forecasting (STLF) is both a planning task and an operational requirement for anticipating high-demand (HD) situations early enough to support preventive action. In practice, distribution forecasting is constrained by the level at which demand can be monitored and acted upon. Privacy and regulatory constraints often prevent the direct operational use of individual smart-meter readings, shifting attention to aggregated signals such as low voltage (LV) feeders, secondary substations (SUBs), and area code (AC) demand [4]–[6]. Aggregation-aware forecasting is therefore essential, since signal behavior, operational relevance, and forecasting difficulty vary across aggregation levels. A key difficulty is that standard evaluation is dominated by normal operating conditions. Since HD intervals are typically sharper and more abrupt than normal demand patterns [2], models can achieve strong global accuracy while still making substantially larger errors around daily peaks [7]. Yet these are exactly the periods that matter most operationally, since peak-load misprediction can have disproportionate effects on cost and reliability [8]. A model that performs well on average but poorly during HD periods is therefore of limited value for distribution-grid operation. STLF is a mature field, and machine learning (ML) methods have long improved over statistical baselines (SBs). In particular, gradient-boosted trees such as XGBoost and LightGBM are widely used because they flexibly exploit calendar, weather, and lag features while remaining scalable and easy to train [9]–[12]. These studies show strong overall accuracy, but mostly under global or average metrics rather than explicit peak-aware evaluation [9]–[12]. Recently, pretrained foundation models (FMs) have emerged as an alternative. Unlike supervised models trained for a specific dataset, they are pretrained on a large collection of time-series data and can be applied in an inferencedriven manner. The Chronos family is a prominent example. Chronos introduced tokenization-based probabilistic forecasting, Chronos-Bolt emphasized faster inference [13], and Chronos-2 [14] extended the framework to a broader multivariate and covariate-aware setting. On electricity datasets
Scaled consumption
HD-threshold
Start/End HD-Period
Isolated HD points
HD-period
Scaled consumption
1
TABLE I DATASET OVERVIEW AND AGGREGATION HIERARCHY ACROSS DISTRIBUTION GRID LEVELS .
0.8
Level
Aggregated from
0.6
AC SUB LV
Multiple SUBs CKW Group (CH) [15] Multiple LV feeders Northern Powergrid (UK) [16] Multiple SMs Northern Powergrid (UK) [16]
0.4
Source
# Entities 115 387 485
0.2
Day 1
Day 2
Day 3
A. Datasets and Aggregation Levels Fig. 1. Typical three-day scaled consumption pattern at SUB level.
such as fev-bench, Chronos-2 attains a 90.7 % win rate over other FMs [14], including TimesFM-2.5, Moirai-2.0, and LagLlama. These properties make Chronos models attractive for operational STLF, especially when quantile outputs and longcontext modeling are useful for peak-aware decision-making. Despite substantial progress in STLF, current evaluation practices remain misaligned with the requirements of distribution-grid operation. Most prior work reports performance averaged over all timestamps, which underrepresents HD intervals where forecast errors are most critical for grid stability, congestion management, and asset utilization. In addition, existing studies often focus on a single aggregation level or a limited number of entities, rather than systematically evaluating DSO-relevant levels such as LV feeders and SUBs. Finally, while recent time-series FMs demonstrate strong general forecasting performance, their effectiveness under peakaware evaluation and their behavior across aggregation levels, particularly with respect to probabilistic outputs, remain insufficiently understood. This work addresses these limitations by investigating whether probabilistic FMs provide a measurable advantage over established statistical and ML approaches for STLF during operationally critical HD intervals across distributiongrid aggregation levels. Specifically, this work makes the following contributions: (i) we introduce a peak-aware evaluation framework that explicitly separates HD and non-HD performance; (ii) we establish a scalable cross-aggregation benchmark across AC, SUB, and LV feeder levels with a large number of entities; and (iii) we analyze the operational behavior of probabilistic FMs, including aggregation-specific quantile selection and runtime considerations. By shifting the evaluation focus from average accuracy to peak-critical performance and by providing actionable guidance on model selection and operating points, this work supports more reliable STLF and improved operational decisionmaking for DSOs. II. DATA AND F ORECASTING S ETUP We study day-ahead STLF across three aggregation levels that are relevant for distribution-grid monitoring and operation: AC, SUB, and LV feeder.
The AC-level data are obtained from the open smart meter (SM) dataset of the Swiss DSO CKW Group [15], which provides aggregated consumption per postal-code region together with the number of contributing SMs. The SUB and LV feeder datasets are taken from the Northern Powergrid open data portal in the United Kingdom [16]. At these three levels, an entity corresponds to one postal-code region, one substation, or one LV feeder, respectively. All datasets cover the period from November 2023 to February 2025. A summary of the datasets used is provided in Table I. B. Preprocessing and High-Demand Definition All time series are processed at the entity level. Duplicate timestamps are removed by aggregating consumption and SM counts per entity and time step. Because the reported number of contributing SMs can vary due to communication or dataquality issues, we use per-meter demand de,t as the primary target signal: Ee,t , de,t = Me,t where Ee,t and Me,t denote aggregated energy consumption and the number of contributing SMs for entity e at time t. The CKW data are available at 15-minute resolution, whereas the Northern Powergrid data are provided at 30minute resolution. To ensure comparability across aggregation levels, all series are aligned to a common 30-minute resolution by consecutive aggregation of CKW intervals. To identify operationally critical periods in a scaleindependent way, we use a rolling HD normalization based on the recent history of each entity, given the lack of publicly available critical threshold data. For entity e and forecast day D, let se,D denote the 99th percentile (p0.99 ) of permeter demand observed during the preceding 14 days. Demand within day D is then normalized as d˜e,t =
de,t , max(se,D , ϵ)
t ∈ D,
where ϵ > 0 ensures numerical stability. A timestamp is classified as HD if d˜e,t ≥ 0.8. All remaining timestamps are treated as non-HD. This definition allows a common peakaware evaluation framework across aggregation levels without requiring explicit operational capacity limits for each entity. Figure 1 shows a typical scaled consumption pattern at the SUB level together with the HD threshold. Where such limits are known in practice, the same framework could be applied directly using those thresholds instead.
C. Forecasting Task and Benchmark Setup
III. F ORECASTING M ETHODS AND E VALUATION We compare statistical baselines, conventional ML models, and probabilistic foundation models. Statistical Baselines: We include two same-time averaging SBs as simple references. One uses the mean at the same time of day over the previous 7 days, and the other over the previous 4 weeks. These baselines are computationally inexpensive and serve as low-complexity reference models. Machine Learning Models: We evaluate LightGBM and XGBoost as representative gradient-boosted tree methods. Both use the feature set described above and produce point forecasts. Gradient boosting remains one of the most widely used and competitive paradigms in applied STLF, combining flexible nonlinear learning with efficient training and robust handling of heterogeneous inputs [9]–[12]. This makes these models important references when testing whether newer FM approaches offer additional benefit when evaluation shifts from average performance to HD behavior. Foundation Models: We further evaluate Chronos-Bolt and Chronos-2 as pretrained probabilistic time-series FMs [13], [14]. Chronos-Bolt emphasizes efficient direct multi-step forecasting and is attractive from an inference-speed perspective. Chronos-2 extends the Chronos line toward broader multivariate and covariate-aware forecasting with longer context support. In contrast to the statistical baselines and standard boosted trees, both models provide quantile forecasts, which enables explicit analysis of conservative versus less conservative operating points under peak-aware evaluation.
Winter
Spring
Jun 24
Sep 24
Summer
500
Peak consumption (Wh)
For each entity, a 1-day-ahead forecast at 30-minute resolution is generated once per day at midnight, resulting in a horizon H of 48 steps. The forecasting objective is evaluated over the period from March 1, 2024 to February 28, 2025 across AC, SUB, and LV feeder data following a rolling-history forecasting protocol. This setup enables a fullyear comparison across seasons: Spring (March-May); Summer (June-August); Autumn (September-November); Winter (December-February), while preserving a sufficient look-back period for both ML and foundation-model forecasting. This results in a large-scale benchmark comprising approximately 41,000 (AC); 141,000 (SUB); and 177,000 (LV feeders) daily forecasts, making it highly scalable and reliable for robust model evaluation across diverse grid conditions. We use a compact set of exogenous covariates consisting of calendar features, holiday indicators, and weather variables. Holiday features are constructed using public holidays in England for Northern Powergrid entities and in Switzerland for CKW entities. Weather variables are obtained via OpenMeteo API [17] and include relative humidity (%), feels-like temperature (°C), dew point temperature (°C), wind speed at 10 m (m/s), and global solar radiation (W/m²). The selection of these weather variables is guided by prior work demonstrating the relevance of meteorological factors for electricity consumption modeling [18], [19].
Autumn
400
300
200
Dec 23
Mar 24
Dec 24
Fig. 2. Daily peak consumption pattern at SUB level.
A. Training and Inference Protocol The forecasting setup follows a rolling-history protocol in which each next-day forecast uses only past information. For the ML models, we use a fixed 90-day seasonal training window rather than daily per-entity retraining, which is computationally infeasible at this scale. As only limited history is available, the closest-matching prior seasonal block was selected using a small validation subset of separate entities from the same datasets that were not part of the final evaluation. This subset showed that warm months exhibit lower demand, whereas cold months show higher and more similar patterns, as illustrated in Figure 2. Accordingly, we forecast Spring 2024 using Winter 2023/24, Summer 2024 using Spring 2024, and both Autumn 2024 and Winter 2024/25 using Winter 2023/24. For the FMs, no retraining is required; instead, each test-day forecast is generated in a rolling manner from recent historical context only, with Chronos-Bolt using up to 2048 timestamps (∼ 42 days) and Chronos-2 using the full 90-day context window under the chosen setup. B. Peak-Aware Evaluation Metric To evaluate both overall forecasting quality and performance during operationally critical periods, all metrics are computed over three timestamp subsets: (i) all points: T , (ii) HD points: THD = {t ∈ T | ỹt ≥ 0.8}, and (iii) non-HD points: TnonHD = {t ∈ T | ỹt < 0.8}. Here, yt and ŷt denote the actual and forecast demand, and ỹt and ỹˆt denote their normalized counterparts. For any subset S ∈ {T , THD , Tnon−HD }, we define the normalized mean absolute error (NMAE) and the mean absolute percentage error (MAPE) as 1 X ỹt − ỹˆt NMAE(S) = |S| t∈S
100 X yt − ŷt MAPE(S) = |S| yt t∈S
Thus, overall, HD, and non-HD metrics are obtained by setting S = T , THD , and TnonHD , respectively. All reported values are computed day-wise and then aggregated over the evaluation period.
TABLE II NMAE & MAPE PERFORMANCE ( MEAN ± STD ) ACROSS AGGREGATION LEVELS FOR ALL MODELS . F OR EACH AGGREGATION LEVEL AND METRIC COLUMN , THE LOWEST MEAN IS SHOWN IN BOLD AND THE SECOND - LOWEST MEAN IS UNDERLINED . Level
Type
Model
NMAE
HD-NMAE
Non HD-NMAE
MAPE
HD-MAPE
Non HD-MAPE
AC
SB SB ML ML FM FM
Last 7-day avg Last 4-week avg LightGBM XGBoost Chronos-Bolt Chronos-2
0.055 ± 0.018 0.066 ± 0.027 0.080 ± 0.032 0.081 ± 0.033 0.051 ± 0.007 0.042 ± 0.007
0.059 ± 0.021 0.073 ± 0.032 0.080 ± 0.027 0.080 ± 0.027 0.053 ± 0.010 0.039 ± 0.007
0.054 ± 0.017 0.064 ± 0.026 0.080 ± 0.033 0.082 ± 0.034 0.051 ± 0.008 0.043 ± 0.008
9.75 ± 3.15 % 11.54 ± 4.79 % 14.65 ± 7.41 % 14.98 ± 7.57 % 9.21 ± 1.57 % 7.68 ± 1.45 %
6.84 ± 2.44 % 8.34 ± 3.61 % 9.11 ± 3.14 % 9.10 ± 3.11 % 6.11 ± 1.12 % 4.53 ± 0.89 %
10.18 ± 3.32 % 11.96 ± 4.95 % 15.49 ± 7.76 % 15.87 ± 7.93 % 9.78 ± 1.65 % 8.25 ± 1.56 %
SUB
SB SB ML ML FM FM
Last 7-day avg Last 4-week avg LightGBM XGBoost Chronos-Bolt Chronos-2
0.074 ± 0.012 0.078 ± 0.011 0.098 ± 0.021 0.098 ± 0.022 0.075 ± 0.010 0.074 ± 0.009
0.117 ± 0.034 0.119 ± 0.031 0.113 ± 0.024 0.113 ± 0.024 0.096 ± 0.021 0.080 ± 0.019
0.068 ± 0.008 0.072 ± 0.010 0.096 ± 0.023 0.096 ± 0.023 0.073 ± 0.009 0.073 ± 0.008
14.96 ± 2.21 % 15.78 ± 2.52 % 20.21 ± 4.62 % 20.42 ± 4.75 % 15.78 ± 2.30 % 15.81 ± 2.25 %
13.54 ± 3.88 % 13.80 ± 3.53 % 13.19 ± 2.74 % 13.20 ± 2.69 % 11.29 ± 2.43 % 9.45 ± 2.21 %
15.01 ± 2.21 % 15.98 ± 2.69 % 21.08 ± 4.93 % 21.31 ± 5.09 % 16.35 ± 2.42 % 16.61 ± 2.36 %
LV
SB SB ML ML FM FM
Last 7-day avg Last 4-week avg LightGBM XGBoost Chronos-Bolt Chronos-2
0.091 ± 0.010 0.096 ± 0.010 0.115 ± 0.018 0.117 ± 0.019 0.094 ± 0.010 0.094 ± 0.009
0.199 ± 0.039 0.197 ± 0.037 0.173 ± 0.031 0.173 ± 0.031 0.160 ± 0.034 0.138 ± 0.030
0.082 ± 0.007 0.088 ± 0.009 0.110 ± 0.019 0.112 ± 0.021 0.089 ± 0.009 0.091 ± 0.008
20.64 ± 2.42 % 21.86 ± 2.82 % 26.96 ± 4.81 % 27.49 ± 5.21 % 22.14 ± 2.76 % 22.80 ± 2.77 %
22.81 ± 4.37 % 22.67 ± 4.16 % 20.07 ± 3.44 % 20.02 ± 3.40 % 18.62 ± 3.75 % 16.14 ± 3.40 %
20.41 ± 2.51 % 21.80 ± 2.99 % 27.52 ± 5.17 % 28.10 ± 5.59 % 22.47 ± 2.88 % 23.36 ± 2.84 %
C. Quantile Selection for Probabilistic Forecasting Chronos-Bolt and Chronos-2 produce multiple quantile forecasts rather than a single point prediction. Since the operating quantile controls the degree of conservativeness, we do not assume that the median forecast is automatically optimal. Instead, we evaluate 11 quantiles, q0.50 , q0.55 , . . . , q0.95 , q0.99 , using the same overall, HD, and non-HD metrics described above. Quantile selection is then performed on a small validation set by jointly considering overall competitiveness and peak-aware performance. The selected quantiles are used in the final benchmark reported in the Results section. IV. R ESULTS A. Performance Across Aggregation Levels Table II summarizes forecasting performance across models and aggregation levels. Chronos-Bolt and Chronos-2 are reported using the aggregation-specific operating quantiles selected on the validation set (Table III); their behavior is analyzed in more detail in Section IV-B. Across all levels, the Chronos models dominate the peak-critical metrics. Chronos2 achieves the best HD-NMAE and HD-MAPE in every setting, and Chronos-Bolt is consistently second-best. This indicates that the FMs are more reliable for capturing peak (HD) intervals while maintaining overall performance that is competitive with the best traditional models. At AC level, this advantage extends to the full evaluation, where Chronos-2 is best on both overall and HD metrics, improving over XGBoost by 48.1 % in overall NMAE and 50.2 % in HD-MAPE. At SUB and LV levels, the Chronos models remain competitive overall, but their clearest advantage appears in the HD performance. At LV level, for example, Chronos-2 is only 3.3 % worse in overall NMAE than the best competing model, yet it improves HD-NMAE by 30.7 % and HD-MAPE by 29.2 %. This pattern is operationally important because
lower aggregation levels exhibit sharper and less regular peaks, making accurate HD forecasting substantially more difficult for both baselines and boosted tree models. The entity-level extremes denoting the best and worst performing entities further confirm the robustness of the Chronos models. Chronos-2 achieves both the lowest best-case and lowest worst-case HD-MAPE across entities, indicating superior consistency and robustness compared to baselines. Its bestcase HD-MAPE reaches 2.47 % at AC, 5.48 % at SUB, and 6.91 % at LV feeder. The robustness is especially evident at SUB and LV feeder levels, where the worst-case HD-MAPE is 39.54 % and 45.67 %, respectively, compared with values above 48 % and 51 % for the SBs and ML models. Seasonal results show the same pattern. At AC and SUB levels, Summer and Spring are the most difficult seasons, respectively, yielding the highest HD-MAPE across model families. Chronos-2 nevertheless shows a much smaller gap between best and worst seasonal HD-MAPE, indicating more stable performance across the year. Its seasonal spread is only 0.68 % points at AC and 1.86 % points at SUB, compared with about 2–3.5 % points for the SBs and ML models. B. Quantile Behavior Quantile selection was performed on a validation set separately for each aggregation level and Chronos model. The rule was: (i) compute, for each candidate quantile, the overall NMAE difference to the best non-Chronos model at the same level and rank candidates in ascending order; (ii) retain the top overall candidates before the HD-NMAE begins to rise again, since HD-NMAE typically decreases as overall NMAE increases up to a turning point; and (iii) among these retained candidates, select the quantile with the lowest HD-NMAE and HD-MAPE. If multiple candidates remained, the one with the smaller overall NMAE gap was chosen. A slightly worse
overall candidate was selected only if it improved both HD metrics relative to the better retained candidate. This rule yielded level-dependent operating quantiles, reported in Table III. At the SUB level, for example, Chronos-2 q0.65 and q0.70 were the two retained candidates because they had the smallest overall NMAE gap to the best non-Chronos model. Here, q0.65 improved on the best non-Chronos model by 6.8 %, whereas q0.70 matched it exactly. Since q0.70 further reduced HD-NMAE by 7.0 % relative to q0.65 , it was selected. The same selection logic was applied at other levels. TABLE III B EST QUANTILE PER MODEL AND AGGREGATION LEVEL . Level
Chronos-Bolt
Chronos-2
AC SUB LV
q0.65 q0.65 q0.65
q0.65 q0.70 q0.70
C. Runtime and Deployment We report median inference latency per 24-hour forecast for Chronos-Bolt and Chronos-2 on CPU and an NVIDIA A100 GPU. Median CPU and GPU inference times are 1.3 s and 0.007 s (Chronos-Bolt) and 9.8 s and 0.2 s (Chronos-2), respectively, resulting in a 185× speedup for Chronos-Bolt and 49× speedup for Chronos-2 on GPU relative to CPU. For deployment, Chronos-2 provides the strongest accuracy–robustness trade-off, while Chronos-Bolt offers competitive accuracy at substantially lower latency. V. C ONCLUSION This paper investigated peak-aware short-term load forecasting across AC, SUB, and LV aggregation levels. Chronosbased FMs consistently outperform established baselines on the operationally critical HD intervals while maintaining competitive overall accuracy. Chronos-2 delivers the strongest performance, reducing HD errors by around 30% relative to the best non-Chronos models without sacrificing overall accuracy. This advantage becomes more pronounced at lower aggregation levels, where load profiles are more volatile and peak behavior is harder to predict. With GPU inference times of 0.007 s for Chronos-Bolt and 0.2 s for Chronos-2 per 24-hour forecast, both models are directly applicable in operational settings. Overall, FMs provide a clear and practical advantage for peak-aware STLF, enabling more reliable monitoring of critical demand periods in distribution grids. Future work will integrate Chronos-2 forecasts into model predictive control to enable uncertainty-aware demand management and gridstability optimization under real-time operational constraints. ACKNOWLEDGMENT This project was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 535389056.
R EFERENCES [1] Department of Transport, UK Gov, “Vehicle licensing statistics: 2024,” https://www.gov.uk/government/statistics/ vehicle-licensing-statistics-2024, 2024, contains public sector information licensed under the Open Government Licence v3.0. [2] M. Savanovic, L. Göberndorfer, and G. Jäger, “Mitigating the charging rush hour,” Heliyon, vol. 10, no. 22, 2024. [3] A. Toleikyte, E. Lecomte, J. Volt, L. Lyons, R. J. C. Roca, A. Georgakaki, S. Letout, A. Mountraki, M. Wegener, A. Schmitz et al., “Clean energy technology observatory: Heat pumps in the european union2024 status report on technology development, trends, value chains and markets,” 2024. [4] European Data Protection Supervisor (EDPS), “EDPS Formal comments on the draft Commission Implementing Regulation on interoperability requirements and non-discriminatory and transparent procedures for access to metering and consumption data,” https://www.edps.europa.eu/system/files/2022-09/22-08-24 access-metering-and-consumption-data en.pdf, 2022, accessed: Nov. 1, 2025. [5] R. Knyrim and G. Trieb, “Smart metering under EU Data Protection Law. International Data Privacy Law,” 2011. [6] D. Lee and D. J. Hess, “Data privacy and residential smart meters: Comparative analysis and harmonization potential,” Utilities Policy, vol. 70, p. 101188, 2021. [7] K. Antoniadou-Plytaria, L. Eriksson, J. Johansson, R. Johnsson, L. Kötz, J. Lamm, E. Lundblad, D. Steen, L. A. Tuan, and O. Carlson, “Effect of short-term and high-resolution load forecasting errors on microgrid operation costs,” in 2022 IEEE PES Innovative Smart Grid Technologies Conference Europe (ISGT-Europe). IEEE, 2022, pp. 1–5. [8] A. Emde, L. Märkle, B. Kratzer, F. Schnell, L. Baur, and A. Sauer, “Effects of load forecast deviation on the specification of energy storage systems,” Designs, vol. 7, no. 5, p. 107, 2023. [9] S. G. K. Uyar, B. K. Ozbay, and B. Dal, “Interpretable building energy performance prediction using XGBoost Quantile Regression,” Energy and Buildings, vol. 340, p. 115815, 2025. [10] M. A. A. Abdalla, A. M. Ishaga, H. A. Osman, M. Elhindi, N. Ibrahim, A. Snani, G. H. A. Hamid, and A. Hammad, “Machine learning-based residential load demand forecasting: Evaluating ELM, XGBoost, RF, and SVM for enhanced energy system and sustainability,” Science in Information Technology Letters, vol. 6, no. 1, pp. 1–15, 2025. [11] H. Musbah and M. Elsaraiti, “Forecasting load consumption: a comprehensive evaluation of deep learning and machine learning techniques,” Electric Power Systems Research, vol. 247, p. 111834, 2025. [12] G. Harikrishnan, T. Premnath, S. Pranav, S. M. Varghese, S. Krishna, and S. Sreedharan, “Machine Learning Approaches for Load Forecasting and Time Series Analysis,” in 2024 7th International Conference on Circuit Power and Computing Technologies (ICCPCT), vol. 1. IEEE, 2024, pp. 1739–1745. [13] A. F. Ansari, L. Stella, C. Turkmen, X. Zhang, P. Mercado, H. Shen, O. Shchur, S. S. Rangapuram, S. P. Arango, S. Kapoor et al., “Chronos: Learning the language of time series,” arXiv preprint arXiv:2403.07815, 2024. [14] A. F. Ansari, O. Shchur, J. Küken, A. Auer, B. Han, P. Mercado, S. S. Rangapuram, H. Shen, L. Stella, X. Zhang et al., “Chronos-2: From univariate to universal forecasting,” arXiv preprint arXiv:2510.15821, 2025. [15] CKW AG, “CKW Open Data Smart Meter: Dataset B - Aggregated smart meter data,” https://open.data.axpo.com/, 2025, accessed: Dec. 6, 2025. [16] Northern Powergrid, “Aggregated smart metering dataset,” https:// northernpowergrid.opendatasoft.com/, 2025, accessed: Nov. 04, 2025. [17] P. Zippenfenig, “Open-meteo.com weather api,” 2023, accessed 2025-12-23. [Online]. Available: https://open-meteo.com/ [18] A. Rahaman, J. Amakor, R. Kazeem, T. Olugasa, O. Ajide, N. Idusuyi, T.-C. Jen, and E. Akinlabi, “Modeling influence of weather variables on energy consumption in an agricultural research institute in ibadan, nigeria,” AIMS energy, vol. 12, no. 1, pp. 256–270, 2024. [19] K. Mosner-Ansong and D. Duah, “The seasonal effects of weather on residential electric-energy usage,” Journal of Energy and Natural Resource Management, vol. 1, no. 1, 2018.