arXiv:2607.04935v1 [cs.DC] 6 Jul 2026
TARE: Tail Aware Evaluation of HPC Job Runtime Prediction Haili Xiao∗
Can Wu∗
Shasha Lu
CNIC, CAS China [email protected]
CNIC, CAS China [email protected]
CNIC, CAS China [email protected]
Xiaoning Wang
Yining Zhao
Rong He
CNIC, CAS China [email protected]
CNIC, CAS China [email protected]
CNIC, CAS China [email protected]
Abstract
ACM Reference Format: Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, and Rong He. 2026. TARE: Tail Aware Evaluation of HPC Job Runtime Prediction. In The 55th International Conference on Parallel Processing (ICPP ’26), September 28–October 1, 2026, Singapore. ACM, New York, NY, USA, 10 pages. https: //doi.org/10.1145/nnnnnnn.nnnnnnn
Runtime estimates affect reservation quality, backfilling opportunities, and queue delay in HPC schedulers. Under heavy tailed workloads, however, averaging over jobs can misrepresent scheduling impact because a small fraction of jobs dominates resource usage. This paper presents an empirical evaluation methodology for HPC job runtime prediction that focuses on the tail, combining GeoAccuracy weighted by resource usage with decile and split analyses. Using production traces from NREL Eagle and ALCF Mira/Intrepid, we compare XGBoost and Last2 against the user provided walltime estimate at submission (UserReq). Across all three datasets, evaluation focused on the tail changes the offline conclusion: MeanAccuracy keeps the methods relatively close, whereas GeoAccuracy reveals clearer separation and makes UserReq’s strength in the upper tail visible. In the top decile, UserReq achieves the highest GeoAccuracy and lowest underestimation rate on all three datasets, and this pattern remains stable across rolling splits. We then translate this signal into a simple hybrid scheduling policy that keeps XGBoost for most jobs and routes the top decile by proxy_cost at submission to UserReq. Online replay on four production queues reduces mean wait time by up to 8% and increases backfilled jobs by 50%–115%. These results show that offline evaluation focused on the tail better characterizes prediction quality relevant to scheduling and informs scheduling policy design.
1
Introduction
In batch scheduled HPC systems, users submit jobs with requested resources such as node counts and walltime limits, and those jobs wait in queue until the scheduler selects them for execution. Under FCFS with EASY backfilling, the scheduler uses a runtime estimate to reserve resources for the job at the head of the queue and to determine whether later jobs can start early without delaying that reservation [13, 16]. Runtime estimation therefore affects reservation quality, backfilling opportunities, and queue performance directly. Overestimation can leave schedulable gaps unused, while underestimation can make planned reservations less reliable. For this reason, HPC job runtime prediction has long been studied as a scheduling problem rather than only as a forecasting problem [25, 30].
GeoAccuracy reveals clearer differences
CCS Concepts • Computer systems organization → Parallel architectures; Resource management; • Computing methodologies → Machine learning. MeanAccuracy
Keywords
GeoAccuracy
Figure 1: Offline evaluation focused on the tail gives new implications for scheduling policy design (see Figure 6).
HPC runtime prediction, performance evaluation, performance metrics, heavy tail workloads, online simulation ∗ Both authors contributed equally to this work as co-first authors.
Because runtime estimates directly affect scheduling decisions, evaluating their quality is itself a scheduling problem. Most runtime prediction studies focus on predictor design and report aggregate accuracy at the job level. But the practical significance of a prediction error depends on which jobs incur it. If a workload has heavy tails, a small fraction of jobs can dominate resource usage. That property is central here: averaging over jobs can obscure the errors with the greatest scheduling consequences. We therefore ask whether weighting accuracy by resource usage reveals predictor differences that job averaged evaluation keeps much closer together. Figure 1
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. ICPP ’26, Singapore © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM https://doi.org/10.1145/nnnnnnn.nnnnnnn 1
ICPP ’26, September 28–October 1, 2026, Singapore
Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, and Rong He
previews this comparison. We compare the weighted geometric metric GeoAccuracy, 𝐴geo , with its arithmetic mean counterpart over jobs, MeanAccuracy, 𝐴mean . We study two runtime predictors and compare them with the user provided walltime estimate at job submission: • XGBoost, a machine learning predictor; • Last2, a heuristic baseline; and • UserReq, the user provided walltime estimate at job submission. We compare these three methods on three production workload datasets: Eagle from NREL and Mira/Intrepid from the Argonne Leadership Computing Facility (ALCF) [18, 22]. All three come from large production HPC systems at national laboratories. Shared preprocessing yields aligned prediction sources. We then apply a consistent offline evaluation across all three datasets and use its findings to design an online replay study on Mira and Intrepid. Before comparing metrics, we verify the workload property that motivates the study in these three traces. A heavy tailed resource usage pattern is not a new discovery of this paper. Prior HPC workload characterization studies have repeatedly reported strong skew and heterogeneity in job size, resource usage, and workload structure on large HPC systems [15, 22, 33]. Our purpose is to show that the same concentration appears in Eagle, Mira, and Intrepid, and to quantify its magnitude. Figure 2 shows that the highest decile alone contributes about 81%–87% of total resource usage across the three datasets. Because HPC resources are scarce and scheduler impact is dominated by large, long running jobs, that concentration makes upper-tail behavior especially important. The CDF view in Figure 2b shows the same concentration from a continuous distribution perspective: Eagle sits farther to the upper left because its jobs consume much smaller node hours, whereas Mira and Intrepid lie farther to the lower right and remain close because the two ALCF traces have similarly large job scales. Under a workload with heavy tails, averages over jobs give every job the same influence, whereas averages weighted by resource usage give more influence to the jobs that consume most node time. As a result, the same method comparison can support different overall conclusions. In our data, 𝐴mean keeps the three methods much closer together, whereas 𝐴geo reveals much clearer differences and makes UserReq’s strength in the upper tail visible on all three datasets. This result is notable because UserReq is the user provided walltime estimate at job submission rather than a learned predictor, but the explanation is that 𝐴geo emphasizes the upper tail of resource usage, where UserReq performs better. This difference also remains stable across most rolling split days, so it is not an artifact of a small sample. The paper follows a three part experimental workflow: shared data preparation, offline evaluation, and online simulation. The offline evaluation weights errors by resource usage and analyzes the upper tail explicitly; the online simulation then tests whether this offline upper-tail signal translates into scheduling outcomes. This emphasis comes from the workload weights and upper-tail analysis, not from any special outlier sensitivity of the geometric mean alone. The contribution is empirical and evaluative rather than algorithmic. We do not propose a new predictor. Instead, this study makes three contributions:
Eagle
Resource usage
80%
Mira
Intrepid
60% 40% 20% 0%
1
2
3
4
5
6
7
Decile
8
9
10
(a) Equal sized decile.
100%
Empirical CDF
80% 60% 40% 20% 0% 10 2
10 1
100
101
102
103
Nodehours (log)
104
105
106
(b) Empirical CDF of node hours.
Figure 2: Evidence of a heavy tail. The highest 10% alone contributes about 81%–87% of total resource usage across the three datasets. (1) We present an empirical evaluation methodology for HPC job runtime prediction that focuses on the tail. It combines GeoAccuracy weighted by resource usage with decile and split level analyses, so offline evaluation reflects the jobs that dominate scheduling impact rather than averages over jobs alone. (2) Using three production workload datasets, we show that evaluation focused on the tail changes the overall comparison of runtime predictors. MeanAccuracy keeps the methods much closer together, whereas GeoAccuracy reveals clearer separation and makes UserReq’s strength in the resource dominant upper tail visible. In the top decile, UserReq achieves the highest GeoAccuracy and the lowest underestimation rate on all three datasets, and this advantage remains stable across rolling split days. (3) We translate this offline signal into a simple hybrid scheduling policy and show in online replay that it improves scheduling performance. The policy keeps XGBoost for most jobs and routes the top decile of jobs by proxy_cost, computed from requested nodes and walltime at submission, to UserReq. Across the short and long production queues on two HPC systems, this policy
2
TARE: Tail Aware Evaluation of HPC Job Runtime Prediction
ICPP ’26, September 28–October 1, 2026, Singapore
reduces mean wait time by up to 8% and increases backfilled jobs by 50%–115%. The rest of the paper is organized as follows. Section 2 presents the methodology, including the offline evaluation metrics and the offline analysis procedure. Section 3 describes the experimental setup, including data preparation, datasets, predictors, and the rolling split offline evaluation procedure. Section 4 presents the offline results, including the overall metric comparison, analysis of the upper tail, and robustness at the split level. Section 5 presents the online simulation, including its design and results. Section 6 reviews related work, and Section 7 concludes the paper.
The log form is what makes the geometric mean natural here. Averaging 𝑒𝑖 summarizes a typical multiplicative miss on an additive scale, and exponentiating back returns the result to the same (0, 1] accuracy scale as 𝐴mean . GeoAccuracy is therefore not an arbitrary alternative summary: it is the consistent aggregate of multiplicative error under weights by resource usage. The emphasis on the workload tail comes from the weights 𝑐𝑖 , not from the use of a geometric mean by itself. GeoAccuracy emphasizes the tail because it weights jobs by resource usage and places most mass on the upper tail of that distribution. It does not make the metric inherently more sensitive to extreme error outliers. This distinction matters because MeanAccuracy and GeoAccuracy answer different questions. MeanAccuracy is an average over jobs. GeoAccuracy is weighted by resource usage and therefore emphasizes the part of the workload whose resource usage most strongly affects scheduling outcomes. To explain the overall metric difference, we also evaluate predictors on equally sized deciles of offline resource usage. Within each dataset, we sort jobs by 𝑐𝑖 and divide them into ten equally sized groups; the tenth group is the top decile. This decile view complements the overall metrics: it explains why GeoAccuracy can reveal clearer differences than MeanAccuracy and localizes the upper tail discussed later in the paper. When the online section later refers to a p90 or top decile routing rule, it uses the same percentile language but applies it to the proxy proxy_cost at submission within each queue-year partition, rather than to the offline resource usage variable 𝑐𝑖 . We also report the underestimation rate 1 ∑︁ UnderRate = 1[𝑦ˆ𝑖 < 𝑦𝑖 ]. (6) 𝑁 𝑖
2 Methodology 2.1 Offline Evaluation Metrics In the batch queue setting, a submitted job carries requested resources and a user walltime estimate, then later yields an observed actual runtime after execution. For job 𝑖, let 𝑦𝑖 denote the actual runtime, 𝑦ˆ𝑖 the predicted runtime, and 𝑐𝑖 = nodes_req × runtime_act denote the resource usage used as the offline weight. We use multiplicative error rather than additive error because HPC job runtimes span a wide range of scales. An additive error measured in seconds is therefore not directly comparable across jobs: the same absolute miss can be minor for a long job but severe for a short job. Taking a log ratio turns multiplicative misses into additive quantities, so ratio errors can be averaged cleanly across jobs. It also treats paired ratio errors symmetrically, so predicting 2𝑦𝑖 and 𝑦𝑖 /2 yields the same penalty. This keeps relative prediction quality separate from the weights by resource usage introduced later in 𝐴geo . The basic log ratio error is 𝑦ˆ𝑖 . (1) 𝑦𝑖 Because the shared preprocessing keeps only jobs with observed runtime above 30 s and the evaluation retains only valid positive predictions, every analyzed row satisfies 𝑦𝑖 > 0 and 𝑦ˆ𝑖 > 0. We therefore do not need a separate branch for zero runtime in the metric definition. The accuracy score for each job is 𝑦ˆ𝑖 𝑦𝑖 , 𝑎𝑖 = exp(−𝑒𝑖 ) = min . (2) 𝑦𝑖 𝑦ˆ𝑖 This yields the familiar arithmetic mean metric over jobs 𝑒𝑖 = log
UnderRate is useful because it exposes a different failure mode from the accuracy metrics. A predictor can achieve competitive MeanAccuracy or GeoAccuracy while still underestimating actual runtime too often. We therefore treat MeanAccuracy, GeoAccuracy, and UnderRate as complementary rather than redundant summaries. Accordingly, the offline evaluation proceeds in three steps. It first compares the overall metrics on the valid rows of each model, then analyzes equally sized deciles of resource usage within each dataset to isolate the highest decile, and finally checks whether the corresponding metric gaps at the split level remain stable across rolling split days. These offline results then motivate the online simulation reported later.
𝑁
𝐴mean =
1 ∑︁ 𝑎𝑖 . 𝑁 𝑖=1
(3)
3
Throughout the paper, we refer to 𝐴mean as MeanAccuracy. To emphasize jobs with high resource usage, we also define the log error weighted by resource usage Í 𝑐 𝑖 𝑒𝑖 . (4) 𝐿usage = Í𝑖 𝑖 𝑐𝑖 The corresponding weighted geometric metric is Ö 𝑐 /Í 𝑐 𝑖 𝑗 𝑗 𝐴geo = exp(−𝐿usage ) = 𝑎𝑖 , (5)
Experimental Setup
Figure 3 summarizes the experimental workflow. This section covers its first two parts: shared data preparation and rolling split offline evaluation on all three datasets. Section 5 then uses the same traces in online replay, which preserves the observed chronology of job arrivals and changes only the planned runtimes seen by the scheduler.
3.1
𝑖
Data Preparation and Date Ranges
We study three production traces from systems at national laboratories: Eagle at NREL and Mira/Intrepid at ALCF [18, 22]. Shared
which is the weighted geometric mean of the job level accuracy scores 𝑎𝑖 . Throughout the paper, we refer to 𝐴geo as GeoAccuracy. 3
ICPP ’26, September 28–October 1, 2026, Singapore
Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, and Rong He
Data Preparation
Offline Evaluation
Online Simulation
Raw job traces Eagle, Mira, Intrepid
Daily rolling splits 100-day train, 1-day test
Queue-year trace replay Mira/Intrepid short and long production queues FCFS + EASY backfilling
Shared preprocessing keep valid terminal states, required attributes, runtime > 30 s
Offline analyses overall metrics, split robustness, equally sized deciles
Aligned prediction sources XGBoost, Last2, UserReq
Offline takeaway upper-tail jobs dominate resource usage; UserReq becomes stronger in the top decile
Hybrid-p90 policy top 10% by proxy_cost → UserReq, otherwise XGBoost
Reported outcomes mean wait time, backfilled-job count
Figure 3: Overview of the experimental workflow. (1) Data preparation produces aligned prediction sources for three traces. (2) Offline evaluation uses daily rolling splits; its finding on the upper tail then informs the hybrid policy used in online simulation. (3) Online simulation replays four production queues from Mira and Intrepid and reports mean wait time and the count of backfilled jobs. data preparation organizes the raw traces around common submit, start, end, and resource request fields; removes invalid or incomplete rows while retaining jobs with valid terminal states; constructs the inputs available at submission that are used by the predictors, including each user’s recent job history needed by Last2; applies the common runtime filter runtime > 30 s; aligns the three prediction sources on consistent jobs; and defines rolling split days after a 100 day warm up window. Their raw submit time spans are 2018-11 to 2023-02 for Eagle, 2014-10 to 2018-12 for Mira, and 2009-12 to 2013-12 for Intrepid. Following the prior Eagle runtime prediction workflow [18], the offline evaluation uses daily rolling splits with a 100 day training window, so evaluated split days begin only after the first 100 days of history. The resulting ranges of split days are 2019-02 to 2023-02 for Eagle, 2015-02 to 2018-12 for Mira, and 2010-03 to 2013-12 for Intrepid. After filtering and alignment of prediction sources, the offline evaluation contains 7,267,468 jobs and 1,435 rolling splits for Eagle, 242,753 jobs and 1,425 splits for Mira, and 271,713 jobs and 1,379 splits for Intrepid.
3.2
a single random holdout. Decile boundaries are computed once per dataset by sorting jobs by resource usage and splitting them into ten equally sized groups; the same boundaries are then applied to all models. At the split level, the analysis unit is each rolling split day rather than individual jobs, and comparisons use the split days where both predictors have valid offline metrics. For each metric, we summarize the fraction of common split days where one predictor outperforms the other: higher values win for metrics such as 𝐴mean and 𝐴geo , while lower values win for metrics such as UnderRate.
4 Offline Results 4.1 Evaluation Focused on the Tail Changes the Offline Conclusion With the heavy tail workload property established in Section 1, Table 1 and Figure 4 report the main overall metrics. The key result is that evaluation focused on the tail changes the offline conclusion. Under 𝐴mean , the methods remain much closer together, whereas 𝐴geo separates them much more clearly. The gap between the maximum and minimum across the three models increases from 0.242 to 0.549 in Eagle, from 0.077 to 0.549 in Mira, and from 0.129 to 0.603 in Intrepid when moving from 𝐴mean to 𝐴geo . This larger separation makes the effect of the workload tail visible at the overall metric level. Under 𝐴mean , the three methods stay close together, especially in Mira and Intrepid. Under 𝐴geo , Last2 drops sharply and UserReq pulls clearly away from the others. The ranking change is therefore a consequence of stronger separation in the tail; the more fundamental point is that evaluation weighted by resource usage reveals much larger performance differences among the methods from a scheduling perspective.
Runtime Predictors and User Estimate
The two runtime predictors and the user provided walltime estimate at job submission play distinct roles. XGBoost is a gradient boosted tree predictor and a strong representative choice for tabular runtime prediction tasks [5, 7, 18]. Last2 is a heuristic baseline that predicts a job as the mean runtime of the same user’s two most recent jobs in the training window, with fallback to the global recent job mean when that user has no history [30]. UserReq uses the submitted walltime request wallclock_req directly.
3.3
Offline Evaluation Procedure
Each rolling split trains on jobs whose end_time lies in the previous 100 days and tests on jobs whose submit_time lies in the next 1 day. For split day 𝑡, the training window is [𝑡 − 100 days, 𝑡) in end_time and the test window is [𝑡, 𝑡 + 1 day] in submit_time. Advancing 𝑡 by one day creates a time ordered sequence of metrics rather than
4.2
Analysis of the Upper Tail Explains Why the Conclusion Changes
Figure 5 shows the decile pattern behind the changed offline conclusion. UserReq is not uniformly strong across the workload. In the 4
TARE: Tail Aware Evaluation of HPC Job Runtime Prediction
ICPP ’26, September 28–October 1, 2026, Singapore
MeanAccuracy (Amean) XGBoost
Metric value
0.8
Last2
GeoAccuracy (Ageo)
UserReq
0.8
0.6
0.6
0.4
0.4
0.2
0.2
0.0
Eagle
Mira
0.0
Intrepid
Eagle
Mira
Intrepid
Figure 4: Overall offline metrics. GeoAccuracy reveals much clearer differences than MeanAccuracy does. lowest deciles of resource usage, it is often the weakest of the three methods because many small jobs are requested conservatively and therefore look inaccurate under a multiplicative metric. This pattern changes sharply as resource usage increases. From the first to the tenth decile of resource usage, UserReq GeoAccuracy rises from 0.009 to 0.501 in Eagle, from 0.071 to 0.774 in Mira, and from 0.167 to 0.738 in Intrepid. By the upper tail, UserReq closes the gap with XGBoost and becomes stronger in the region with the highest resource usage: the tenth decile in Eagle,
the seventh through tenth deciles in Mira, and the ninth and tenth deciles in Intrepid. This decile shift explains why GeoAccuracy and MeanAccuracy support different offline conclusions. MeanAccuracy gives equal weight to the many jobs with low resource usage where UserReq is weak. GeoAccuracy gives much more weight to the upper deciles, where UserReq becomes progressively more accurate and where resource usage is concentrated. The overall metric difference is therefore driven by a clear change in predictor behavior across deciles.
Table 1: Overall offline metrics.
Dataset Eagle
XGBoost Last2 UserReq Mira XGBoost Last2 UserReq Intrepid XGBoost Last2 UserReq
GeoAccuracy in decile
0.511 0.440 0.269 0.671 0.617 0.594 0.660 0.531 0.545
0.529 0.101 0.650 0.609 0.252 0.801 0.608 0.162 0.765
Eagle
100%
4.3
𝐴mean 𝐴geo UnderRate
Model
Top Decile Results Localize the Advantage Relevant to Scheduling
Figure 5 already shows the decile trend. The highest decile localizes the advantage relevant to scheduling because it is the part of the workload that carries most of the resource usage. Table 2 reports the exact metrics for the top decile. In every dataset, UserReq achieves the highest GeoAccuracy in the top decile and the lowest underestimation rate. Relative to XGBoost, its GeoAccuracy in the top decile rises from 0.470 to 0.501 in Eagle, from 0.640 to 0.774 in Mira, and from 0.643 to 0.738 in Intrepid. Over the same datasets, UnderRate falls from 72.7% to 34.1%, from 77.9%
39.9% 52.5% 11.0% 54.0% 50.3% 17.1% 48.7% 55.8% 11.5%
Mira
100%
80%
80%
80%
60%
60%
60%
40%
40%
40%
20%
20%
20%
0%
1
2
3
4 5 6 7 8 Equally sized decile of resource usage
9
10
0%
1
2
3
4 5 6 7 8 Equally sized decile of resource usage XGBoost
Last2
Intrepid
100%
9
10
0%
1
2
3
4 5 6 7 8 Equally sized decile of resource usage
9
10
UserReq
Figure 5: Why and where UserReq becomes a useful estimate. UserReq is weak on many jobs with low resource usage but improves sharply toward the higher deciles, eventually becoming competitive with or stronger than XGBoost and Last2 in the upper tail. 5
ICPP ’26, September 28–October 1, 2026, Singapore
Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, and Rong He
Table 2: Top decile offline metrics. Jobs are split into ten equally sized deciles by offline resource usage; “Top decile usage” reports the share of total usage carried by the highest decile in each dataset. Dataset
Model
Eagle
XGBoost Last2 UserReq XGBoost Last2 UserReq XGBoost Last2 UserReq
Mira
Intrepid
Top-decile usage
Top-decile GeoAccuracy
Top-decile UnderRate
86.5%
0.470 0.100 0.501 0.640 0.317 0.774 0.643 0.201 0.738
72.7% 82.5% 34.1% 77.9% 74.0% 32.2% 75.7% 80.3% 26.0%
82.3%
81.4%
to 32.2%, and from 75.7% to 26.0%, which corresponds to reductions of 38.6, 45.7, and 49.7 percentage points, respectively. It is also the only method that performs best on both GeoAccuracy and UnderRate in the top decile on all three datasets. Across datasets, XGBoost remains competitive on GeoAccuracy in the tail, but its underestimation rate stays much higher; Last2 is weaker on both tail metrics. This pattern across datasets is already clear from the reported values. Taken together, these top decile results show a consistent pattern in the upper tail with clear implications for scheduling. The same prediction source that looks weak on many small jobs becomes both more accurate and less prone to underestimation on the jobs that dominate resource usage. That is the direct offline signal behind the hybrid policy studied later: if one prediction source is simultaneously safer and more accurate in the upper tail that matters most for scheduling, then routing that tail differently becomes a natural online design choice.
4.4
98.9% on Mira, and 99.6% on Intrepid. Variation at the split level is still expected because each day has a different test set size and a different upper tail composition. Even so, the dominant direction is consistent across the rolling sequence. The clearer differentiation under GeoAccuracy, together with UserReq’s much lower underestimation rate, is therefore a stable effect at the workload level rather than an artifact of a few extreme days.
5
5.1
proxy_cost = nodes_req × wallclock_req.
Following the protocol at the split level in Section 3, we treat each rolling split day as one analysis unit and compare XGBoost with UserReq only on the split days where both predictors have valid offline metrics. Table 3 reports the resulting win rates across split days. The signal at the split level remains stable under both metrics, but its direction depends on the metric: 𝐴mean usually favors XGBoost, whereas 𝐴geo usually favors UserReq. XGBoost has the higher 𝐴mean on 67.3%–87.4% of split days, while UserReq has the higher 𝐴geo on 68.9%–81.0% of split days. The reversal at the split level is therefore not coming from a few isolated days; it is the dominant direction across the rolling sequence.
UserReq higher 𝐴geo
UserReq lower UnderRate
Eagle Mira Intrepid
87.3% 67.3% 87.4%
68.9% 81.0% 79.1%
97.4% 98.9% 99.6%
(7)
Under Hybrid-p90, jobs above a p90 cutoff for each queue and year in proxy_cost go to UserReq, and all others stay on XGBoost. This policy does not assume that UserReq should replace machine learning predictors everywhere. It tests whether routing the tail of jobs with high proxy_cost to UserReq improves scheduling outcomes in online simulation. The p90 choice is intended as a simple, interpretable operating point motivated by the offline results rather than a claim of global optimality. Figure 6 summarizes the routing rule.
Table 3: Robustness at the split level, XGBoost vs. UserReq. XGBoost higher 𝐴mean
Hybrid Scheduling Policy
The hybrid policy applies UserReq only to jobs with high values of a proxy for resource usage available at submission. We use
Results at the Split Level Show That the Signal Is Stable
Dataset
Online Simulation
The offline results show that UserReq becomes stronger precisely in the part of the workload that dominates resource usage and therefore matters most for scheduling. This section turns that finding into an online test: Section 5.1 defines the hybrid scheduling policy, Section 5.2 describes the replay setup, and Section 5.3 reports the resulting behavior at the queue level.
UserReq’s advantage on underestimation is even more stable. It achieves the lower UnderRate in 97.4% of split days on Eagle, 6
TARE: Tail Aware Evaluation of HPC Job Runtime Prediction
ICPP ’26, September 28–October 1, 2026, Singapore
Table 4: Precision, recall, and recall weighted by cost for each queue when the proxy_cost top decile is used to identify the top decile of real cost.
Submitted job nodes_req, wallclock_req
Compute proxy available at submission proxy_cost = nodes_req × wallclock_req
Dataset
Queue
Top decile by proxy_cost?
Mira Mira Intrepid Intrepid
prod-short prod-long prod-short prod-long
Yes
Recall
Cost-weighted Recall
72.9% 57.5% 61.1% 69.5%
81.4% 91.2% 80.4% 88.9%
89.8% 93.5% 88.5% 92.4%
No
Yes: route to UserReq planned runtime = wallclock_req
summary reports the relative change in mean wait time and the count of backfilled jobs across four queues drawn from the long and short production queues on Mira and Intrepid. We focus on these four queues because short and long production queues are common and representative in HPC operations, and together they expose both short job and long job scheduling regimes on two different systems. Across these four replay queues, proxy_cost aligns well with the heavy tail region. At the same queue and year granularity used by the hybrid policy, the proxy_cost top decile attains 80.4%–91.2% recall and 88.5%–93.5% recall weighted by cost with respect to the top decile of real cost. Table 4 summarizes the values for each queue.
No: keep XGBoost planned runtime = runtime_pred
Scheduler uses the selected planned runtime for reservation and backfilling
Figure 6: The Hybrid-p90 scheduling policy used in online simulation. Jobs above a p90 cutoff for each queue and year in proxy_cost are routed to UserReq; the rest remain on XGBoost.
5.2
Precision
5.3
Online Simulation Setup
Hybrid Replay Validates the Offline Signal
We test whether the offline signal carries through to outcomes at the queue level on Mira and Intrepid. The hybrid rule routes only jobs with high proxy_cost to UserReq. In the replay, Hybrid-p90 sends the top decile of jobs by proxy_cost within each partition for a queue and year to UserReq and leaves the rest on XGBoost. Figure 7 summarizes the comparison across cases in a single view. For each of four queues, the diamond and its adjacent label mark the average relative change across years under (Hybrid-p90 − XGBoost)/XGBoost, the pale background bar shows the yearly span, and the smaller points show the individual cases for each queue and year. The yearly cases span 2015–2018 for Mira and 2010–2013 for Intrepid. Hybrid-p90 improves both metrics in all four cases. In every case, it reduces mean wait time relative to XGBoost and increases the number of backfilled jobs. The clearest gains in wait time appear in the short queues: Intrepid/prod-short shows an average relative wait reduction of 8.4%, and Mira/prod-short shows a reduction of 7.8%. Intrepid/prod-long still improves by 2.8%, while Mira/prod-long is nearly neutral at 0.3% lower average wait because one later year offsets earlier gains. The backfilling side is even more uniformly favorable. The gains in short queues are especially large: 115.3% in Mira/prod-short and 87.8% in Intrepid/prod-short. The gains in long queues are also substantial. They reach 77.5% in Intrepid/prod-long and 50.1% in Mira/prod-long. These percentages should be interpreted directionally rather than as throughput multipliers, because the baseline number of backfilled jobs is small in some long queue years. Even with that caveat, the figure shows a clear and repeatable effect at the queue level: routing based on proxy_cost tends to preserve more backfill opportunities while also shortening wait time. This pattern suggests that the short production queues leave more room
To test the implications of the offline results for scheduling, we supplement the offline evaluation with online replay on Mira and Intrepid. We replay each queue and year partition independently, for example Mira/prod-short/2016. In this replay, the simulator preserves the original job chronology in timestamp order and changes only the planned runtime presented to the scheduler. Each job keeps its trace submit time and requested node count. The predictor under test supplies the planned runtime seen by the scheduler. For XGBoost, we use the same hyperparameter setting as in the original runtime prediction workflow of Menear et al. [18]. Reservation and backfilling decisions use that planned runtime, while node release uses the actual runtime from the trace. The online simulation is intentionally simple. Each partition for a queue and year is simulated independently under FCFS with EASY backfilling. At each job submission or completion event, the scheduler first computes the earliest feasible reservation for the job at the head of the queue. It then scans later queued jobs in submit order and starts any job whose requested nodes and predicted runtime fit without delaying that reservation. Once started, a job occupies its requested nodes for its actual runtime. Those nodes are returned to the partition when the job completes. Partition capacity is inferred from the observed peak concurrent node usage in the trace. To keep the comparison fair, the online simulation is restricted to jobs for which XGBoost, Last2, and UserReq all produce valid predictions after runtime filtering and prediction validity checks. The simulator reports mean wait time and the count of backfilled jobs. For the online simulation comparison, we evaluate the hybrid routing policy defined above rather than treating the offline result as only a comparison at the model level. The online simulation 7
ICPP ’26, September 28–October 1, 2026, Singapore
Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, and Rong He
Mean waittime
Backfilled jobs
Intrepid / prodshort 8.4%
87.8%
7.8%
115.3%
Queue
Mira / prodshort
Intrepid / prodlong 2.8%
77.5% 0.3%
50.1%
Mira / prodlong 20.0%
15.0%
10.0%
5.0%
0.0%
0%
50%
100%
150%
200%
Improvement: (Hybridp90 XGBoost) / XGBoost Figure 7: Hybrid replay improves outcomes at the queue level across years. Each row is one queue. The pale background bar shows the span of yearly relative changes, the small points show individual years, and the diamond with its adjacent label marks the average for each queue under (Hybrid-p90 − XGBoost)/XGBoost. Left: lower mean waittime is better. Right: more backfilled jobs are better. for backfill gains in our replay, a result that is consistent with prior studies that characterize workloads [22]. The points for each year in Figure 7 show that the average diamonds are not driven by a single anomalous year. The short queues on both systems, together with Intrepid/prod-long, improve in every simulated year on both mean wait time and the count of backfilled jobs. The only mixed case is Mira/prod-long: 2016 and 2017 improve on both metrics, 2018 is slightly worse on both, and 2015 is effectively tied because both policies backfill the same small number of jobs. The dot range summary therefore preserves the yearly pattern directly instead of hiding instability behind a single average. These results link the offline tail effect to scheduling outcomes. A simple policy at submission time that invokes UserReq only on jobs with high proxy_cost improves both mean wait time and backfilling in all four queues. The simulation is intentionally limited: it uses partitions for each queue and year with FCFS scheduling and EASY backfilling rather than a production scheduler deployment. We therefore interpret the result as evidence of scheduling impact rather than full scheduler validation.
6
directly [6, 34]. Recent work has focused largely on improving the predictor itself through ensemble methods, deep learning, feature engineering, or domain specific and two step estimation strategies [4, 7, 8, 18, 21]. Among these, tree ensemble methods such as XGBoost [5] are especially common because they handle heterogeneous submission features and nonlinear interactions well, and they have become strong practical baselines in HPC runtime prediction studies [7, 18]. In this literature, performance is usually summarized with aggregate arithmetic mean metrics such as RMSE, MAE, or overall accuracy like scores, sometimes together with risk measures focused on underestimation [12]. A related but distinct line of work studies predictive uncertainty through conditional quantiles and prediction intervals. Large scale quantile regression makes quantile estimation practical on very large datasets [35], and conformalized quantile regression adds prediction intervals that are valid for finite samples [23]. These methods target uncertainty quantification, not evaluation oriented toward schedulers under resource usage with heavy tails. Our question is different from both predictor construction and interval estimation. We reevaluate existing prediction sources under different aggregation rules. Under resource usage with heavy tails, the choice of metric can reveal clearer differences among predictors and therefore support different conclusions about their quality for scheduling. The focus is therefore on predictor evaluation rather than predictor construction.
Related Work
We place this paper in the context of three adjacent areas: runtime prediction methods, studies of user provided walltime estimates and backfilling, and work that evaluates prediction quality through scheduling outcomes. This paper is closest to the third area because it asks how runtime predictors should be evaluated, rather than proposing a new predictor.
6.1
6.2
User Provided Runtime Estimates and Backfilling
User provided runtime estimates have been studied for decades because they affect reservation length, backfilling opportunities, and queue performance directly. Much of this literature emphasizes that user estimates are systematically overestimated or otherwise inaccurate in aggregate [2, 29]. However, the same line of work also shows that their operational value should not be dismissed: under
HPC Runtime Prediction
HPC job runtime prediction has a long history spanning models based on historical information, predictors based on logs, and more recent machine learning systems. Early work showed that historical behavior can support scheduling decisions [24], and later studies used system logs or job attributes to predict completion time more 8
TARE: Tail Aware Evaluation of HPC Job Runtime Prediction
ICPP ’26, September 28–October 1, 2026, Singapore
EASY backfilling, conservative user estimates can preserve reservation safety and may even help utilization or queue behavior in some settings [13, 19, 28, 31]. Subsequent work studied adjusted user estimates, controls for underestimation, and scheduler mechanisms that replace or augment user estimates with predictions generated by the system [9, 10, 12, 14, 27, 30]. Runtime prediction has also been linked to queue time prediction and scheduler performance more broadly [3, 17, 25, 32]. Our results complement this literature. We do not claim that user estimates are the most useful prediction source across the whole workload. Instead, we show that they become especially valuable on the upper tail of resource usage and that this effect becomes visible only under evaluation focused on the tail.
6.3
policy design under HPC workloads with heavy tails. They also suggest that no single prediction source serves the entire workload equally well: machine learning predictors remain useful on the bulk of jobs, while the user provided walltime estimate at job submission carries valuable signal in the upper tail. Online schedulers should therefore combine multiple prediction signals through a hybrid policy rather than forcing one source to serve the whole workload.
References [1] Omar Aaziz, Jonathan Cook, and Mohammed Tanash. 2018. Modeling Expected Application Runtime for Characterizing and Assessing Job Performance. In 2018 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, Belfast, 543–551. doi:10.1109/CLUSTER.2018.00070 [2] Cynthia Bailey Lee, Yael Schwartzman, Jennifer Hardy, and Allan Snavely. 2005. Are User Runtime Estimates Inherently Inaccurate? In Job Scheduling Strategies for Parallel Processing, David Hutchison, Takeo Kanade, Josef Kittler, Jon M. Kleinberg, Friedemann Mattern, John C. Mitchell, Moni Naor, Oscar Nierstrasz, C. Pandu Rangan, Bernhard Steffen, Madhu Sudan, Demetri Terzopoulos, Dough Tygar, Moshe Y. Vardi, Gerhard Weikum, Dror G. Feitelson, Larry Rudolph, and Uwe Schwiegelshohn (Eds.). Vol. 3277. JSSPP ’05, Berlin, Heidelberg, 253–263. doi:10.1007/11407522_14 [3] Nick Brown, Gordon Gibb, Evgenij Belikov, and Rupert Nash. 2024. Predicting Accurate Batch Queue Wait Times on Production Supercomputers by Combining Machine Learning Techniques. Concurrency and Computation: Practice and Experience 36, 15 (July 2024), e8112. doi:10.1002/cpe.8112 [4] Fengxian Chen. 2023. Job Runtime Prediction of HPC Cluster Based on PCTransformer. Journal of Supercomputing 79, 17 (Nov. 2023), 20208–20234. doi:10. 1007/s11227-023-05470-2 [5] Tianqi Chen and Carlos Guestrin. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, San Francisco, CA, USA, 785–794. doi:10.1145/2939672.2939785 [6] Xin Chen, Charng-Da Lu, and Karthik Pattabiraman. 2013. Predicting Job Completion Times Using System Logs in Supercomputing Clusters. In 2013 43rd Annual IEEE/IFIP Conference on Dependable Systems and Networks Workshop (DSN-W). IEEE, Budapest, Hungary, 1–8. doi:10.1109/DSNW.2013.6615513 [7] Xiaomeng Chen, Hui Zhang, Hanli Bai, Chunming Yang, Xujian Zhao, and Bo Li. 2020. Runtime Prediction of High-Performance Computing Jobs Based on Ensemble Learning. In Proceedings of the 2020 4th International Conference on High Performance Compilation, Computing and Communications. ACM, Guangzhou China, 56–62. doi:10.1145/3407947.3407968 [8] Hyunjoon Cheon, Jinseung Ryu, Jaecheol Ryou, Chan Yeol Park, and Yo-Sub Han. 2023. ARED: Automata-Based Runtime Estimation for Distributed Systems Using Deep Learning. Cluster Computing 26, 5 (Oct. 2023), 2629–2641. doi:10.1007/ s10586-021-03272-w [9] Václav Chlumský and Dalibor Klusáček. 2022. Improving Accuracy of Walltime Estimates in PBS Professional Using Soft Walltimes. In Job Scheduling Strategies for Parallel Processing, Dalibor Klusáček, Corbalán Julita, and Gonzalo P. Rodrigo (Eds.). Vol. 13592. JSSPP ’22, Cham, 192–210. doi:10.1007/978-3-031-22698-4_10 [10] Hang Cui, Keichi Takahashi, Yoichi Shimomura, and Hiroyuki Takizawa. 2025. Clustering Based Job Runtime Prediction for Backfilling Using Classification. In Job Scheduling Strategies for Parallel Processing, Dalibor Klusáček, Julita Corbalán, and Gonzalo P. Rodrigo (Eds.). Vol. 14591. JSSPP ’24, Cham, 40–59. doi:10.1007/ 978-3-031-74430-3_3 [11] A.B. Downey. 1997. Predicting Queue Times on Space-Sharing Parallel Computers. In Proceedings 11th International Parallel Processing Symposium. IEEE Comput. Soc. Press, Genva, Switzerland, 209–218. doi:10.1109/IPPS.1997.580894 [12] Yuping Fan, Paul Rich, William E. Allcock, Michael E. Papka, and Zhiling Lan. 2017. Trade-Off Between Prediction Accuracy and Underestimation Rate in Job Runtime Estimates. In 2017 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, Honolulu, HI, USA, 530–540. doi:10.1109/CLUSTER.2017.11 [13] D.G. Feitelson and A.M. Weil. 1998. Utilization and Predictability in Scheduling the IBM SP2 with Backfilling. In Proceedings of the First Merged International Parallel Processing Symposium and Symposium on Parallel and Distributed Processing. IEEE Comput. Soc, Orlando, FL, USA, 542–546. doi:10.1109/IPPS.1998.669970 [14] Eric Gaussier, David Glesser, Valentin Reis, and Denis Trystram. 2015. Improving Backfilling by Using Machine Learning to Predict Running Times. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. ACM, Austin Texas, 1–10. doi:10.1145/2807591.2807646 [15] Jie Li, George Michelogiannakis, Brandon Cook, Dulanya Cooray, and Yong Chen. 2023. Analyzing Resource Utilization in an HPC System: A Case Study of NERSC Perlmutter. arXiv:2301.05145 [cs] doi:10.48550/arXiv.2301.05145 [16] David A. Lifka. 1995. The ANL/IBM SP Scheduling System. In Job Scheduling Strategies for Parallel Processing, Gerhard Goos, Juris Hartmanis, Jan Leeuwen,
Evaluation Metrics and Scheduling Impact
Prior work has argued that prediction quality should ultimately be judged by its effect on scheduling rather than by offline accuracy alone [11, 25, 26]. Trace studies also show that large HPC workloads are heterogeneous and that aggregate averages can hide important structural differences across jobs, users, and resource classes [1, 20, 22]. At the same time, most runtime prediction papers still treat the evaluation metric as fixed and place the novelty in the predictor. Relative to that evaluation oriented line of work, this paper combines three elements. First, it uses an overall metric, 𝐴geo , that weights error by resource usage alongside the metric based on averaging over jobs, 𝐴mean . Second, it uses equally sized deciles of resource usage to explain why GeoAccuracy reveals clearer differences among the methods. Third, it uses online replay to test whether those offline differences translate into scheduling outcomes. We are not aware of prior HPC runtime prediction work that centers this combination of evaluation focused on the tail, explanation at the decile level, and validation at the queue level, while explicitly linking the offline metric choice to the upper tail of the workload that matters most for scheduling.
7
Conclusion
This paper makes three contributions. First, it presents an empirical evaluation methodology for HPC job runtime prediction that focuses on the tail and combines GeoAccuracy weighted by resource usage with analyses at the decile and split levels. Second, it shows on three production HPC workloads that evaluation focused on the tail changes the offline conclusion. 𝐴mean keeps the compared methods relatively close, whereas 𝐴geo reveals clearer separation and makes UserReq’s strength in the resource dominant upper tail visible. In the top decile, UserReq achieves the highest GeoAccuracy and the lowest underestimation rate on all three datasets, and this advantage remains stable across rolling split days. Third, it translates this offline signal into a simple hybrid scheduling policy and shows in online replay that it improves scheduling performance. This policy uses proxy_cost at submission as the routing rule. It keeps XGBoost for most jobs and routes the top decile to UserReq. Across the short and long production queues on two HPC systems, it reduces mean wait time by up to 8% and increases backfilled jobs by 50%–115%. These results indicate that offline evaluation focused on the tail should inform the assessment of runtime predictors and scheduling 9
ICPP ’26, September 28–October 1, 2026, Singapore
Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, and Rong He
Dror G. Feitelson, and Larry Rudolph (Eds.). Vol. 949. JSSPP ’95, Berlin, Heidelberg, 295–303. doi:10.1007/3-540-60153-8_35 [17] Kevin Menear, Kadidia Konate, Kristi Potter, and Dmitry Duplyakin. 2024. Tandem Predictions for HPC Jobs. In Practice and Experience in Advanced Research Computing 2024: Human Powered Computing. ACM, Providence RI USA, 1–9. doi:10.1145/3626203.3670547 [18] Kevin Menear, Ambarish Nag, Jordan Perr-Sauer, Monte Lunacek, Kristi Potter, and Dmitry Duplyakin. 2023. Mastering HPC Runtime Prediction: From Observing Patterns to a Methodological Approach. In Practice and Experience in Advanced Research Computing. ACM, Portland OR USA, 75–85. doi:10.1145/3569951.3593598 [19] A.W. Mu’alem and D.G. Feitelson. 2001. Utilization, Predictability, Workloads, and User Runtime Estimates in Scheduling the IBM SP2 with Backfilling. TPDS ’01 12, 6 (June 2001), 529–543. doi:10.1109/71.932708 [20] Alessio Netti, Woong Shin, Michael Ott, Torsten Wilde, and Natalie Bates. 2021. A Conceptual Framework for HPC Operational Data Analytics. In 2021 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, Portland, OR, USA, 596–603. doi:10.1109/Cluster48925.2021.00086 [21] Alan L. Nunes, Bernardo Gallo, Bruno Lopes, Felipe A. Portella, José Viterbo, Lúcia M. A. Drummond, Luciano Andrade, Miguel de Lima, Paulo J. B. Estrela, and Renzo Q. Malini. 2025. Two-Step Estimation Strategy for Predicting Petroleum Reservoir Simulation Jobs Runtime on an HPC Cluster. CPE 37, 4-5 (2025), e70026. doi:10.1002/cpe.70026 [22] Tirthak Patel, Zhengchun Liu, Raj Kettimuthu, Paul Rich, William Allcock, and Devesh Tiwari. 2020. Job Characteristics on Large-Scale Systems: Long-Term Analysis, Quantification, and Implications. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, Atlanta, GA, USA, 1–17. doi:10.1109/SC41405.2020.00088 [23] Yaniv Romano, Evan Patterson, and Emmanuel J. Candes. 2019. Conformalized Quantile Regression. In Advances in Neural Information Processing Systems 32. Curran Associates, Inc., Red Hook, NY, USA, 3538–3548. [24] Warren Smith, Ian Foster, and Valerie Taylor. 1998. Predicting Application Run Times Using Historical Information. In Job Scheduling Strategies for Parallel Processing, Gerhard Goos, Juris Hartmanis, Jan Van Leeuwen, Dror G. Feitelson, and Larry Rudolph (Eds.). Vol. 1459. JSSPP ’98, Berlin, Heidelberg, 122–142. doi:10.1007/BFb0053984 [25] Warren Smith, Valerie Taylor, and Ian Foster. 1999. Using Run-Time Predictions to Estimate Queue Wait Times and Improve Scheduler Performance. In Job Scheduling Strategies for Parallel Processing, Gerhard Goos, Juris Hartmanis, Jan Van Leeuwen, Dror G. Feitelson, and Larry Rudolph (Eds.). Vol. 1659. JSSPP ’99, Berlin, Heidelberg, 202–219. doi:10.1007/3-540-47954-6_11 [26] Ozan Sonmez, Nezih Yigitbasi, Alexandru Iosup, and Dick Epema. 2009. TraceBased Evaluation of Job Runtime and Queue Wait Time Predictions in Grids. In Proceedings of the 18th ACM International Symposium on High Performance Distributed Computing. ACM, Garching Germany, 111–120. doi:10.1145/1551609. 1551632 [27] Wei Tang, Narayan Desai, Daniel Buettner, and Zhiling Lan. 2010. Analyzing and Adjusting User Runtime Estimates to Improve Job Scheduling on the Blue Gene/P. In 2010 IEEE International Symposium on Parallel & Distributed Processing (IPDPS). IEEE, Atlanta, GA, USA, 1–11. doi:10.1109/IPDPS.2010.5470474 [28] Dan Tsafrir. 2010. Using Inaccurate Estimates Accurately. In Job Scheduling Strategies for Parallel Processing, Eitan Frachtenberg and Uwe Schwiegelshohn (Eds.). Vol. 6253. JSSPP ’10, Berlin, Heidelberg, 208–221. doi:10.1007/978-3-64216505-4_12 [29] Dan Tsafrir, Yoav Etsion, and Dror G. Feitelson. 2005. Modeling User Runtime Estimates. In Job Scheduling Strategies for Parallel Processing, David Hutchison, Takeo Kanade, Josef Kittler, Jon M. Kleinberg, Friedemann Mattern, John C. Mitchell, Moni Naor, Oscar Nierstrasz, C. Pandu Rangan, Bernhard Steffen, Madhu Sudan, Demetri Terzopoulos, Dough Tygar, Moshe Y. Vardi, Gerhard Weikum, Dror Feitelson, Eitan Frachtenberg, Larry Rudolph, and Uwe Schwiegelshohn (Eds.). Vol. 3834. JSSPP ’05, Berlin, Heidelberg, 1–35. doi:10.1007/11605300_1 [30] Dan Tsafrir, Yoav Etsion, and Dror G. Feitelson. 2007. Backfilling Using SystemGenerated Predictions Rather than User Runtime Estimates. TPDS ’07 18, 6 (June 2007), 789–803. doi:10.1109/TPDS.2007.70606 [31] Dan Tsafrir and Dror G. Feitelson. 2006. The Dynamics of Backfilling: Solving the Mystery of Why Increased Inaccuracy May Help. In 2006 IEEE International Symposium on Workload Characterization. IEEE, San Jose, CA, 131–141. doi:10. 1109/IISWC.2006.302737 [32] Chiara Vercellino, Alberto Scionti, Giuseppe Varavallo, Paolo Viviani, Giacomo Vitali, and Olivier Terzo. 2023. A Machine Learning Approach for an HPC Use Case: The Jobs Queuing Time Prediction. Future Generation Computer Systems 143 (June 2023), 215–230. doi:10.1016/j.future.2023.01.020 [33] Qiqi Wang, Yu Shen, and Jing Li. 2021. User-Level Workload Analysis for Supercomputers. In 2021 The 4th International Conference on Software Engineering and Information Management. ACM, Yokohama Japan, 68–73. doi:10.1145/3451471. 3451483 [34] Qiqi Wang, Hongjie Zhang, Jing Li, Yu Shen, and Xiaohui Liu. 2022. Predicting Job Finish Time Based on Parameter Features and Running Logs in Supercomputing System. The Journal of Supercomputing 78, 17 (Nov. 2022), 18551–18577. doi:10.
1007/s11227-022-04582-5 [35] Jiyan Yang, Xiangrui Meng, and Michael W. Mahoney. 2013. Quantile Regression for Large-Scale Applications. In Proceedings of the 30th International Conference on Machine Learning, Vol. 28. PMLR, Atlanta, Georgia, USA, 881–887.
10