SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances Taeyoon Kima , Kyumin Kima , Kyunghwan Kimb , Hayoung Kima , Seungwoo Jeonga , Moohyun Songc , Kyungyong Leea,∗ a Hanyang University, Department of Data Science, Seoul, 04763, Republic of Korea b KT, Seoul, 06763, Republic of Korea c Hanyang University, Department of Artificial Intelligence, Seoul, 04763, Republic of Korea
arXiv:2604.24548v1 [cs.DC] 27 Apr 2026
Abstract Cloud vendors offer discounted spot instances to maximize surplus resource utilization, but these instances are subject to the risk of sudden interruption. Traditional pricing datasets have been employed to predict this risk, yet recent policy changes by cloud vendors have diminished their effectiveness. To promote spot instance usage, public cloud vendors provide instant availability datasets to help users mitigate interruption risks. While existing research utilizing this data has proposed methods to reduce interruptions, these studies have primarily focused on single-node instances, overlooking the stability of multi-node environments widely adopted for modern cloud workloads. This paper proposes SpotVista, a system that recommends a resource pool of reliable and cost-efficient multi-node spot instances by leveraging various publicly available datasets. To achieve this, SpotVista collects a large-scale multi-node availability dataset while overcoming significant query limitations. Through a thorough analysis of multi-node spot instance availability behavior, SpotVista establishes a methodology for recommending cost-efficient and reliable multi-node configurations. To evaluate how effectively the proposed methodology reflects multi-node availability and cost efficiency, extensive real-world interruption experiments were conducted. The results demonstrate that SpotVista outperforms the state-of-the-art work, SpotVerse, achieving 81.28% greater availability and 2.84% more cost savings in a multi-region setup. When compared to a publicly available service, AWS SpotFleet, SpotVista provides 21.6% higher stability and 26.3% greater cost savings. Keywords: Spot Instances, Availability Score, Cloud Computing
1. Introduction Cloud computing has fundamentally altered the way of consuming computing resources by offering elastic capacity without the overhead of managing physical infrastructure. To provide users with such flexibility, cloud service vendors maintain resources beyond peak demand, which results in a surplus. To maximize the utilization of these otherwise idle resources, cloud vendors offer them at significant discounts of up to 90%, referring to them as spot instances. This service model has been adopted by most major cloud vendors, including AWS, Azure, GCP, Alibaba, and IBM. Initially, the price of spot instances was dynamically adjusted based on supply and demand, and an instance may be reclaimed, an event known as spot instance interruption, when the market price exceeds a user’s bid price. Instance interruptions pose a significant threat to application reliability, requiring users to prepare a mechanism to deal with such events. To assist users in this endeavor, cloud vendors have historically provided public datasets, most notably spot pricing data. This dataset was extensively used in prior research to analyze price behavior with the interruption events [1, 2, 3, 4, 5], enhance application reliability [6, 7, 8, 9, 10, 11, 12, 13], and predict optimal bid prices ∗ Corresponding author
Email address: [email protected] (Kyungyong Lee) Preprint submitted to Elsevier
to minimize interruption risk [14, 15, 16, 17, 18, 19]. However, recent changes in cloud vendor policies have weakened the correlation between spot prices and instance terminations, reducing the effectiveness of these price-based predictive models [20, 21]. In response to this shift, cloud vendors have introduced new availability-focused datasets, such as AWS and Azure Spot Placement Score (SPS), which provide real-time indicators of spot instance availability [22, 23]. While services like SpotLake have emerged to archive and provide access to these datasets [24, 25], they are limited in scope. Specifically, they offer availability information only for single-node instances and do not adequately address the requirements of modern distributed applications, such as large-scale machine learning model training and big data processing, which often depend on multi-node environments. Cheon et al. [26] presented the discrepancy of a single node SPS dataset when inferring availability of multi-node spot instances pool, and they emphasized the inadequacy of single-node metrics for provisioning reliable, large-scale resource pools. Furthermore, the dataset archive service provides only raw dataset, and users must choose cost-efficient and reliable spot instances manually that can be a very challenging task. Although cloud vendors offer native recommendation services like AWS SpotFleet, their internal operational mechanisms are not disclosed, and the effectiveness of their recommendations has not been independently validated. April 28, 2026
Number of requested instances
To address these limitations by exploring the availability characteristics of multi-node spot instance environments, we propose a heuristic to efficiently collect a comprehensive multi-node availability dataset followed by performing thorough analysis for the dataset. Based on this dataset, this work introduces a novel availability score to quantify the stability of multi-node spot instance configurations. This scoring mechanism serves as the foundation for a recommendation algorithm designed to identify spot instance pools that achieve an optimal balance between cost-efficiency and high availability. The effectiveness of the proposed system is validated through extensive experiments on real-world cloud infrastructure, involving 59,560 spot requests across 127 unique instance types. The proposed algorithm demonstrates 81.28% higher availability and 2.84% greater cost savings than the state-of-the-art system, SpotVerse [27], and provides 21.6% higher availability with 26.3% more cost savings compared to the commercial AWS SpotFleet service. The collected multi-node availability dataset and the recommendation engine are made publicly available as a web service to benefit the research community and cloud practitioners1 . The main contributions of this paper are as follows: • The development of a novel sampling heuristic for the efficient collection of a comprehensive multi-node spot instance availability dataset, addressing significant query limitations. • A detailed analysis of the collected dataset, revealing key temporal and spatial characteristics of multi-node spot instance availability. • The design and validation of a spot instance recommendation engine that quantifies both availability and costefficiency, whose effectiveness is demonstrated through extensive real-world experiments. • The deployment of a publicly available web service that provides both the collected multi-node availability data and the recommendation engine.
1 72.5% 5 43.8% 10 25.0% 15 10.0% 20 25 3.8% 30 2.5% 35 2.5% 40 2.5% 45 2.5% 50 0.0% 0 20 40 60 80 Successful Instance Types (%)
93.8%
100
Figure 1: Success rate of spot instance allocation for types with a single-node SPS of 3. Each type was requested 50 times, and the bars indicate the fraction of types that achieved at least the corresponding number of successful allocations.
In response, cloud vendors have shifted to providing datasets directly related to spot instance availability, and AWS and Azure have started to offer metrics such as interruption ratios over the past month and real-time availability data, exemplified by SPS [22, 23]. Although the internal calculation details are not disclosed, these metrics are described as indicators of immediate spot instance availability. For example, AWS’s SPS assigns a score of 1 (Low), 2 (Medium), or 3 (High) to each instance type, where a higher score signifies greater availability. While users can access these datasets through management consoles or APIs, they are often constrained by query limitations. To overcome these constraints and facilitate easier data access, a web service named SpotLake has been developed [24, 25]. However, a critical limitation of existing availability datasets and services is their focus on single-node instances that are misaligned with the demands of modern cloud applications. Workloads such as large-scale deep learning model training [31] and distributed big data processing [32] inherently require multinode environments and incur significant computational costs. Consequently, leveraging spot instances to mitigate these costs has become a compelling strategy for distributed systems, leading to a substantial body of research. This research spans various domains, including DNN training [33, 34], big data processing [35], and general-purpose web services [9, 11, 13]. The majority of this existing work concentrates on reactively handling interruptions after an instance has been reclaimed. Some studies have explored proactively selecting low-risk instances to prevent interruptions. For example, Kim et al. proposed a method to enhance stability by predicting SPS values [36], and SpotVerse [27] references SPS and interruption frequency score to infer spot instance availability. Yet, these approaches, like others, are constrained by their reliance on single-node spot instance data, limiting their applicability to distributed environments. For large-scale multi-node deployments, an availability metric that reflects the stability of an entire batch of many instances is essential. Cheon et al. [26] examined the inadequacy of single-node SPS for predicting availability in multi-instance requests. Their experiments showed that while a single spot instance request
2. Spot Instance and Availability For spot instance users, balancing cost savings with operational stability is a primary consideration, given the dynamically changing prices and the inherent risk of interruption events. To assist users, public cloud vendors offer various datasets that are related to spot instance prices and stability. In the past, particularly within auction-based models (e.g., AWS [28], Azure [29], Alibaba [30]), these datasets were widely employed to estimate interruption risks [14, 15]. For instance, AWS provided regular updates and historical archives of spot prices, which facilitated a significant body of research on their stable usage [1, 3, 16]. However, recent policy changes have diminished the correlation between spot prices and interruptions. The reduced frequency of price updates has consequently limited the effectiveness of the previous research [21, 20].
1 https://spotvista.ddps.cloud
2
succeeded in all trials, the success rate decreased substantially as the requested count increased. For a batch of 50 instances, the allocation rate dropped to 20%, indicating that single-node SPS is not a reliable metric for multi-node deployments. Figure 1 further illustrates this discrepancy. Each instance type with a single-node SPS of 3 was requested 50 times. The fraction of types achieving successful allocation decreases sharply as the required number of nodes in the horizontal axis grows. Fewer than half of the types succeed when 10 or more instances are requested, and no type achieved full allocation at the maximum request of 50 instances. This behavior arises because instances of the same type within an Availability Zone (AZ) are provisioned from a shared capacity pool [37]. The single-node SPS reflects only whether at least one instance can be allocated from this pool, and does not capture how many instances the pool can simultaneously satisfy. As a result, a high single-node SPS does not guarantee the availability of a large multi-node request, since the request may exceed the remaining capacity of the shared pool. This gap between single-node and multi-node availability cannot be inferred from single-node SPS alone, motivating the need for a dataset that directly characterizes multi-node availability.
control over query costs by adjusting both the query interval p and the step size T s . However, this approach introduces a trade-off where a specific node count is re-evaluated only after min a delay of (⌊ TmaxT−T ⌋ + 1) × p minutes. During this period, any s SPS fluctuations could lead to data staleness, necessitating an analysis of this proposed method’s data integrity. 3.1.1. Assessing Sampled Dataset Integrity A key parameter in USQS is the step size, T s , which determines the granularity of the sampling. A larger T s reduces query overhead but increases the risk of missing the exact transition points. Considering both query overhead and the correctness of the dataset, a step size of T s = 5 is adopted, which provides a favorable balance between capturing significant availability trends and maintaining low query overhead. To quantitatively assess the potential information loss from this sampling strategy with the given T s , entropy, a measure of uncertainty from information theory, is employed [38]. The entropy H(X) is defined as X H(X) := − p(x) log p(x) (1) x∈X
where X is a discrete random variable representing the SPS value observed at a given query point, X denotes the set of all possible SPS outcomes (i.e., {1, 2, 3}), and p(x) is the probability of observing SPS value x. Intuitively, entropy quantifies the degree of unpredictability in SPS transitions: when all outcomes are equally likely, entropy reaches its maximum, indicating maximum uncertainty; conversely, when the distribution is heavily concentrated on a single outcome, entropy approaches zero, reflecting highly predictable behavior. A higher entropy value signifies greater randomness in the data, which increases the risk of missing critical changes when sampling periodically. Conversely, data with lower entropy exhibits more predictable patterns, making it well-suited for a sampling-based collection approach. With T s = 5 and T max = 50, there are 11 possible discrete numbers of nodes, {1, 5, 10, . . . , 50}. If the distribution of these outcomes were uni1 form, the probability of each would be 11 , yielding a maximum possible entropy of 3.4594 bits. However, the real SPS value exhibits non-uniform, skewed distributions [24]. To measure the actual entropy, multi-node SPS data was collected from 844 instance types between January 26th and 29th, 2025, and the collected dataset yielded a measured entropy of 2.5052 bits. This value is significantly lower than the theoretical maximum for a uniform distribution. The lower entropy confirms that the temporal behavior of score change patterns is not random but contains predictable patterns. This result validates that the USQS method can effectively capture the state of multi-node availability with minimal information loss while substantially reducing query overhead, a claim further substantiated in the evaluation section.
3. Collecting a Batch of Spot Instances Availability Dataset AWS offers SPS values for multiple instances of up to 50, but its query limitations make it difficult to retrieve this data efficiently. For example, within a 24-hour window, only 50 distinct query scenarios are allowed, and queries for the same configuration with different node counts are treated as separate requests. This restriction results in significant query overhead when gathering data for all combinations of instance types, regions, and node counts. Even using the optimal query distribution proposed by SpotLake [24], approximately 3,300 queries are needed for all instance types for a single-node with 66 separate accounts. To query more than a single-node, the query overhead and required accounts increase proportionally. For example, querying SPS values from one to 50 nodes would require 165,000 queries, spread across 3,300 accounts. Moreover, to maintain dataset timely, these queries must be run periodically, making it nearly impossible to query all possible combinations. To address this, we propose a sampling heuristic to query the dataset in a timely manner. 3.1. Uniform Spacing Query Sampling To mitigate the high query overhead required to track multinode SPS values, we propose a heuristic named Uniform Spacing Query Sampling (USQS). Instead of querying the entire range of node counts in every collection cycle, USQS probes a single target node count, T c , at each periodic interval of p minutes. The target count is systematically alternated across cycles, incremented by a fixed step size, T s . Specifically, after querying for T c nodes, the subsequent query targets T c + T s nodes. This process continues until the target count exceeds the predefined maximum, T max , at which point it resets to the minimum, T min . This method allows for granular
3.2. Tracking Score Transition Point While the USQS method effectively reduces query overhead, its sampling nature implies it may not identify the exact node 3
count at which an SPS score transition occurs. For scenarios where higher precision is critical, this section introduces an alternative approach, Tracking Score Transition Points (TSTP), designed to locate the precise transition points without the potential for sampling-induced data loss. This method is based on the observation that the SPS for a given instance type within an AZ exhibits a monotonically nonincreasing property. As the number of requested nodes increases, the SPS value either remains the same or decreases. The TSTP approach leverages this behavior to efficiently and accurately pinpoint the specific node counts where the SPS value changes from 3 to 2, or 2 to 1. For a given range of node counts [T min , T max ], two key transition points are defined: • T 3: The largest node count for which the SPS is 3. • T 2: The largest node count for which the SPS is 2. By definition, T 3 ≤ T 2. These two values concisely represent the multi-node availability metric; any request for a number of nodes up to T 3 will have an SPS of 3, while requests between T 3+1 and T 2 will have an SPS of 2, and so on. A standard binary search algorithm can identify these points within O(log(T max − T min )) API queries, avoiding the need to query every possible node count.
instances based on criteria such as instance family, category, or region. Each candidate node is then evaluated using a cost score and an availability score, which serve as the basis for recommendations. 4.1. Quantifying Spot Instance Cost To establish a quantitative metric for comparing spot instance costs, the total cost to fulfill a user’s request with a specific instance type must first be calculated. Let pi be the unit price for an instance type i with CPUi cores. For a resource request defined by a total number of CPU cores, RC , the required number of instances, Ni , is calculated as Ni = ⌈RC /CPUi ⌉. Consequently, the total cost for a pool composed of instance type i is ci = pi × N i . Let C be the set of total costs for all candidate instances. To enable a fair comparison, these costs are normalized to a score, CS i , on a scale of [0.0, 100.0] using Equation 2. CS i = 100 ×
Cmin Ci
(2)
In the equation, Cmin denotes the minimum of the cost set C. The scoring function computes the ratio of the minimum cost to the cost of instance i, producing a score of 100 for the cheapest instance and proportionally lower scores for more expensive ones. The score directly reflects how cost-efficient an instance is relative to the best available option, without requiring any distributional assumptions. A key advantage of this inverse min-scaling formulation is its independence from the overall cost distribution. Unlike MinMax scaling [39], which compresses scores when extreme outliers exist and exaggerates marginal differences when all costs are close, the proposed metric evaluates each instance solely against the minimum cost. This provides stable and interpretable scores regardless of the cost distribution’s shape.
Reducing Query Overhead via Caching and Early Stopping. To reduce the query overhead of a standard binary search, two optimization techniques are applied. First, since SPS values for an instance type tend not to fluctuate drastically over short periods [24], the T 3 and T 2 values from the previous data collection cycle are cached. In the next cycle, the search begins near the cached value rather than at the midpoint of the entire [T min , T max ] range, allowing the algorithm to narrow the search range with a single API call. Second, an early stopping mechanism terminates the binary search when the search range (T high − T low ) becomes smaller than a predefined threshold e, as an approximate transition point within a small error margin is sufficient for assessing instance stability. These two optimizations are complementary: caching accelerates the early stages of the search by leveraging temporal continuity, while early stopping eliminates the diminishing-return queries in the final stages where additional query yields only a marginal approximate error reduction.
4.2. Quantifying Spot Instance Availability To quantitatively assess the availability of a spot instance, a composite scoring model is proposed. This model is based on the time-series analysis of T 3, defined as the maximum number of nodes for which the SPS score remains at 3. The availability score, AS i , for a given instance type i is derived from three key characteristics of its T 3i time-series data including magnitude, trend, and volatility. First, the overall magnitude of availability is captured by the area under the T 3i time-series curve over the observation period, normalized to a scale of [0.0, 1.0], denoted A3i , using a MinMax scaler across all candidate instances. A higher A3i indicates a greater and more consistent capacity for multi-node deployments. Second, the trend of availability over time, denoted by mi , provides a predictive measure of future stability. An increasing T 3 trend is a positive indicator, while a decreasing trend signals potential risk. The trend is calculated as the normalized slope of a first-order linear regression model fitted to the T 3i data [40]. Third, the volatility of availability, represented by σi , quantifies the stability of the instance. High fluctuation of T 3 is
4. Recommendation Engine for Multi-Node Spot Instances While real-time availability datasets from cloud vendors assist in the initial selection of spot instances, this information alone is insufficient for identifying configurations that are both costefficient and reliable over time. Raw availability data does not capture historical stability patterns or provide a comparative cost analysis across diverse instance types. To address this gap, this section details a recommendation engine that systematically identifies optimal multi-node spot instance pools by leveraging historical multi-node availability and price datasets. The recommendation process begins with user-specified requirements for a computing resource pool, primarily the total number of CPU cores (RC ) or the total memory size (R M ). Users can also apply optional filters to narrow the set of candidate 4
that meet the user’s initial requirements. These scores are then combined into a final score, S i , using a tunable weight parameter, W ∈ [0.0, 1.0], as shown in Equation 4.
T3
50 40 30 20 10 0
Thu.
Sat.
Mon. Time
Wed.
T3
(a) consistently high T 3 (score : 100) 50 40 30 20 10 0
Thu.
Sat.
Mon. Time
Wed.
(c) positive slope T 3 (score : 59)
50 40 30 20 10 0
Thu.
Sat.
Mon. Time
Wed.
S i = W × AS i + (1.0 − W) × CS i
(b) consistently low T 3 (score : 0)
T3
T3
50 40 30 20 10 0
Thu.
Sat.
Mon. Time
The weight W allows users to prioritize availability over cost, or vice versa. Empirical analysis, detailed in the evaluation section, indicates that a weight of W = 0.5 provides a robust balance between the two objectives, and it is used as the system’s default setting. Simply selecting the single instance type with the highest score (S i ) is often suboptimal in a multiple spot instances resource pool. Fulfilling a large resource request with numerous instances of a single instance type makes the entire workload vulnerable to an interruption event due to the specific type. To mitigate this risk, the proposed system constructs a heterogeneous pool of recommended instance types. Given the scores for all candidate instances, the goal is to select a subset of instance types and determine the number of nodes for each, forming a pool that maximizes the total quality score while satisfying the user’s resource requirements. This problem can be formulated as an Integer Linear Programming (ILP) problem, which is known to be NP-hard [41]. While ILP solvers can identify globally optimal solutions for the score maximization objective, this formulation alone does not address the need for instance type diversity. Relying on a small number of high-scoring types makes the resulting pool vulnerable to correlated interruption events, which motivates the explicit consideration of diversity during pool construction. Incorporating diversity into the ILP formulation, however, is structurally difficult. The number of selected types is not a linear function of the allocation variables, and expressing this requirement requires additional binary indicator variables with linking constraints. The resulting search space grows rapidly as the number of candidate instances increases, which makes ILPbased formulations impractical for real-time recommendation in large candidate spaces, where a quantitative comparison is presented in Section 6. To address this issue, SpotVista introduces a greedy heuristic that enhances the diversity of the selected instance types while minimizing the degradation of the aggregated total scores. The detailed procedure is presented in Algorithm 1. The algorithm begins by sorting all candidate instances in descending order of their final scores with initialization of variables (Lines 5-8), including the recommendation pool P and the previous allocation for the top-ranked instance, x prev top . The core of the algorithm is a main loop that iterates through the sorted candidates (Line 9). In each iteration, the next highestscoring instance type is added to the pool P to increase diversity (Line 10). Subsequently, the system recalculates the node allocation for every instance currently in the pool (Lines 13-17). This is achieved by distributing the total required resources, Rreq , among the pool members proportionally to their individual scores (Line 14) and then calculating the necessary number of nodes for each (Line 15). After each reallocation, the algorithm evaluates two termination conditions (Line 18). The first, Xcurr [C sorted [0]] ≥ x prev top ,
Wed.
(d) periodic changing T 3 (score : 45)
Figure 2: Example of quantifying spot instance availability score by using the area, slope, and standard deviation of SPS score change over time
undesirable for reliable workloads. This metric is calculated as the normalized standard deviation of the T 3i values. A larger σi corresponds to greater instability. These three components are combined to compute the final availability score, AS i , as shown in Equation 3. AS i = 100 × (A3i × (1.0 + λ × (mi − σi )))
(4)
(3)
The formula begins with the base magnitude score A3i and adjusts it through a scaling coefficient λ that bounds the maximum influence of the trend and volatility components. Both mi and σi are normalized to [0.0, 1.0], so the adjustment term (mi − σi ) is bounded within [−1.0, 1.0]. The coefficient λ therefore determines the maximum percentage by which the base score can be adjusted. Extensive empirical evaluation reveals that λ being 0.1 shows the best performance for spot instance reliability modeling, under which the formula applies a bonus of up to 10% for a positive trend and a penalty of up to 10% for high volatility. The justification for this choice is presented in the evaluation section. Figure 2 provides examples of how the availability score (AS i ) quantifies different temporal patterns in T 3 values. The horizontal axis in each plot represents time, and the vertical axis represents the T 3 value, capped at a maximum of 50 in this scenario. An instance with a consistently high T 3 value (Figure 2a) achieves a perfect score of 100. This is the result of a maximal area component (A3i ), combined with zero penalty for volatility (σi = 0) and no adjustment for trend (mi = 0). Conversely, a consistently low T 3 (Figure 2b) yields a score of 0 because its base area score is zero. Dynamic patterns are scored based on a combination of factors. The instance in Figure 2c displays a positive trend. Although it incurs a minor penalty for its volatility (σi > 0), it receives a substantial bonus for the upward slope (mi > 0), resulting in a favorable score. In contrast, the instance with periodic fluctuations (Figure 2d) is heavily penalized for its high volatility. With a slope that converges to zero (mi ≈ 0), it receives no trend-based bonus, leading to a comparatively low score. These examples demonstrate how the AS i score holistically evaluates availability by integrating the magnitude, trend, and stability of an instance’s T 3 profile. 4.3. Spot Instance Recommendation The recommendation engine first calculates the availability score (AS i ) and the cost score (CS i ) for all candidate instances 5
Algorithm 1 Greedy Heuristic for Spot Instance Pool Formation
SpotVista: Multiple Spot Instances Recommendation System Serverless Spot Instances Recommendation Service
1: procedure FormHeterogeneousPool(C, Rreq ) 2: Input: C: Set of candidate instances with scores S i . 3: Input: Rreq : Total required resources (e.g., CPU cores). Output: Xbest : A map of instance types and their count. 4: 5: 6: 7: 8: 9: 10: 11: 12: 13: 14:
Serverless Backend
Get Latest Data
Object Storage Latest Data
Send Recommendation Results
C sorted ← Sort C by score S i in descending order P ← ∅ ▷ The set of instance types in the current pool Xbest ← ∅ x prev top ← ∞ ▷ Allocation of the top-ranked instance for i ∈ C sorted do P ← P ∪ {i} ▷ Add next best instance to the pool P S total ← j∈P S j Xcurr ← new map for all j ∈ P do S ▷ Score-based allocation R j ← S totalj × Rreq
.csv
.csv
Price 1-50 SPS Query
API Endpoint
Query & Scoring Function
Spot Instance Data Collector
Time Series Database Historical Raw Data Store Historical SPS Data
.csv
Y/M/D/h-m-s Round Robin
[ 1, 5, 10, ⋯, 40, 45, 50 ]
Users Every 10 Minute Rule
Uniform Spacing Query Sampling
SPS Collector
Upload Raw Data
Price Collector
Figure 3: Implementation of the proposed system
R
x j ← ⌈ CPUj j ⌉ 16: Xcurr [ j] ← x j 17: end for 18: if Xcurr [C sorted [0]] ≥ x prev top or Xcurr [i] = 0 then 19: return Xbest 20: end if 21: Xbest ← Xcurr 22: x prev top ← Xcurr [C sorted [0]] 23: end for 24: return Xbest 25: end procedure 15:
Show Latest Price, T3 and T2 on The Web Page
Static Web
recommendation module. The files are saved with a Year-MonthDate-Time naming convention to include historical data. The recommendation module is designed using a serverless architecture [42] to flexibly handle irregular recommendation requests from users. When users provide their resource requirement, such as minimum number of CPU cores or memory size, along with optional information such as specific instance types, regions, or the maximum number of returned instance types, this data is sent to an API Endpoint service, which acts as the endpoint for the recommendation service. It forwards the request to a Function-as-a-Service platform, which filters the instance types that meet the user’s requirements and retrieves the historical T 3 values of the target instances via a serverless time-series database service. The recommendation score is calculated using this information and then returned to the user. The web service itself is served as static HTML files via object storage, and the instance recommendation functionality interacts with the backend RESTful API through the Fetch API. The monthly infrastructure cost of operating SpotVista is approximately $55 USD, comprising $30 for an instance running the data collection pipeline, $20 for the time-series database service, and $5 for object storage. The recommendation module incurs no additional fixed cost, as it is deployed on a serverless architecture that charges only per invocation. This low operational overhead demonstrates the practicality of continuously maintaining a large-scale multi-node availability dataset as a public service.
detects that the top-ranked instance’s allocation has stopped decreasing, meaning the newly added instance’s score is too low to meaningfully redistribute resources away from the dominant type, and further additions of even lower-scored instances would yield diminishing diversification. The second, Xcurr [i] = 0, indicates that the latest addition receives no allocation under the score-proportional distribution (Line 14), so it would contribute no capacity to the pool. When either condition is met, the algorithm returns the previous iteration’s allocation, Xbest (Line 19), preserving the last state in which diversification was effective. If the algorithm does not terminate, it updates Xbest with the current allocation and saves the new top-ranked node count to x prev top for the next iteration’s comparison (Lines 21-22). This process ensures that the recommendation is iteratively refined to balance score quality with instance diversity.
5. SpotVista Implementation
6. Evaluations
SpotVista is publicly available as a web-service. The architecture of the implemented system targeting spot instances is shown in Figure 3. The Data Collector component at the bottom applies the proposed USQS heuristic for SPS data collection where queries are periodically generated every 10 minutes with the target number of nodes ranging from 5 to 50, incremented by 5. For price data, we directly utilize the price dataset provided by cloud vendors. The collected dataset is stored in an object storage, available for direct access by users or for use by the
We analyze the efficiency and operational overhead of the multi-node SPS data collection module. We also demonstrate the effectiveness of the instance recommendation algorithm through extensive experiments in real-world spot instance environments. An intuitive method to assess spot instance stability and interruption probability is to continuously run spot instances and record any interruptions [36, 43, 44]. Such an approach is valid for small-scale experiments. However, the cost can increase significantly as the scale of the experiment grows. To control the 6
7.5 5.0 2.5 ry + rm he Unifaocing Bienaarch CBaScheBS + CyacStop p S l r S + Ea (a) Different query heuristics comparison
50
40
40
30
30
20
20
10
10
0
USQS 10
20
30
40
50
0
Sequential queries per period
(b) Overhead and error rate of USQS
3.0
Score Difference (%)
10.0
50
Number of Queries
12.5
Difference from Real Values
Num. Queries Difference
Number of Queries
Difference from Real Values
30 25 20 15 10 5 0
High Mid Low
2.5 2.0 1.5 1.0 0.5 0.0
5 10 15 20 25 30 35 40 45 50
Target Capacity
(c) Error ratio analysis
Mean Absolute Error
Figure 4: Multi-node SPS dataset collection overhead and completeness of the proposed sampling query heuristic
cost of interruption experiments, Wu et al. [45] demonstrated that continuously running spot instances is not necessary. Instead, periodically sending spot instance requests and observing whether the requests succeed can effectively model spot instance stability. Li et al. [46] adopted the same probing-based methodology to characterize regional spot availability. It issues lightweight launch requests that immediately terminate upon success, thereby measuring availability at low cost. Spot-andScoot [47] further corroborated this finding. They reported that spot request outcomes rarely overestimate actual capacity and that interruptions of co-located instances of the same type exhibit strong temporal correlation. Given the large-scale experiments involving many instance types in this paper, we adopt this methodology. We periodically sent spot requests, recorded success or failure, and generated interruption data based on these observations. Using the result, we aim to answer the following research questions. • RQ-1 Does the sampling-based multi-node SPS collection heuristic effectively capture key information without significant loss? • RQ-2 What characteristics are observed in the multi-node SPS dataset? • RQ-3 Does the proposed availability scoring mechanism work as expected? • RQ-4 How does the spot instance recommendation algorithm perform compared to the state-of-the-art approaches?
7 6 5 4 3 2 1 0 0
5 10 15 20 25 30 35 40 45 50 Step Size (Ts)
Figure 5: Mean Absolute Error of the USQS heuristic as a function of step size
error of 0.9. USQS requires only a single query per cycle; its median error remains low, though the 100-minute re-query interval can cause larger deviations for highly dynamic instances. This marginal precision loss is a tolerable trade-off for the substantial reduction in query overhead. Query Overhead vs. Error Rate. Figure 4b compares USQS against sequential scanning with increasing query counts (10–50 per cycle). As the number of queries increases linearly, the T 3 error decreases only marginally, confirming that the 10×–50× overhead reduction of USQS far outweighs its minimal precision loss. Impact on Collected SPS Data. Figure 4c shows the percentage difference in average SPS values between USQS and the Full Scan, categorized by historical SPS volatility (High, Mid, Low). Regardless of volatility, the maximum deviation remains below 3%, confirming that the USQS sampling has a negligible effect on the integrity of the collected time-series data used by the availability scoring model.
6.1. Effectiveness of the SPS Query Heuristics To answer RQ-1, this section evaluates the trade-off between query overhead and data integrity for the proposed heuristics. Figure 4 presents this analysis using a dataset collected from October 12–15, 2024, for 15 instance types across 7 regions. The ground truth was established by a Full Scan that queries the SPS of all 50 target node counts every 10 minutes.
Sensitivity to Step Size. To further validate the choice of T s = 5, we simulated the USQS heuristic over the ground-truth dataset for all step sizes from T s = 1 to T s = 50. Figure 5 shows that the MAE follows a U-shaped curve driven by two competing error sources: for small T s , the prolonged round-robin cycle causes temporal staleness (e.g., 500 minutes per cycle at T s = 1), while for large T s , wide spacing between probes misses SPS transition points. The MAE reaches its minimum region at T s = 3–5, where values remain below 2.1. Among these, T s = 5 is selected
Analysis of Different Query Heuristics. Figure 4a compares the T 3 estimation error (primary vertical axis) and query overhead (secondary vertical axis) across four strategies in the horizontal axis. Plain Binary Search (BS) achieves minimal error but requires 12 queries per cycle. Adding caching and early stopping (e=4) reduces this to about 7 queries with a negligible average 7
use1
usw2
apne1
8/5
8/6
8/7 8/8 DateTime
euc1
apse1
Seasonal Variation
26 24 22 20 18 16 14
az6
az4
az2
az3
8/5
8/6
8/7 8/8 DateTime
az1
az5
T3
euw1
T3
22 20 18 16 14 12 10
8/4
1.0
8/9
8/10
8/11
(a) T 3 change patterns across different AWS region (2025)
8/9
8/10
8/11
(b) T 3 change patterns by AZ in the AWS us-east-1 region (2025)
AWS
0.5
8/4 Azure
0.0 0.5 2025-04
2025-05
2025-06
2025-07
2025-08
2025-09
2025-10
2025-11
2025-12
(c) Weekly seasonal component of T 3 extracted via MSTL decomposition for AWS and Azure
Figure 6: Spatial, temporal, and seasonal characteristics of multi-node T 3 availability, combining short-term observations with long-term MSTL decomposition [48].
Table 1: Temporal stability analysis of T 3 seasonal patterns
Metric
AWS
Azure
MSTL Variance Decomposition Daily (24h) 0.783 Weekly (168h) 0.039 Trend 0.060 Residual 0.003
0.030 0.041 1.115 0.029
Seasonal Strength (FS ) Daily Weekly
0.997 0.934
0.510 0.608
Bai-Perron Amplitude Stability Daily breakpoints 4 Daily max variation ±9% Weekly breakpoints 3 Weekly max variation ±7%
5 ±44% 4 ±28%
business hours, consistent with prior spot availability studies [49, 36]. The same pattern holds at the AZ-level within us-east-1. Both levels also reveal that certain locations consistently offer higher multi-node availability, confirming that location selection is a critical factor for reliable spot deployments. Long-term Seasonal Stability. To verify whether these patterns persist over longer horizons and across vendors, MSTL decomposition [48] is applied to extended datasets from AWS (April– November 2025; 1,114 types, 17 regions) and Azure (same period; 1,705 types, 67 regions). Two metrics quantify the stability pattern; the seasonal strength FS [50], which measures how strongly a periodic component dominates over residual noise ([0.0, 1.0]), and the Bai-Perron structural break test [51], which detects statistically significant shifts in seasonal amplitude. Figure 6c visualizes the extracted weekly seasonal component, and Table 1 summarizes the quantitative results. For AWS, the daily cycle dominates the overall variance, both daily and weekly FS exceed 0.93, and the Bai-Perron test shows amplitude variation of at most ±9%. These results confirm that the cyclical behavior observed in the short-term analysis persists over eight months, providing an empirical basis for the historical pattern-based availability scoring proposed in this work. It is noticeable that Azure exhibits a structurally different pattern. The trend component dominates variance rather than seasonal cycles, FS values are substantially lower (0.51 daily, 0.61 weekly), and amplitude shifts reach ±44%. Directly applying the proposed scoring methodology to Azure would therefore require adaptation mechanisms to account for the weaker and less stable seasonality, which we further discuss in Section 8.
because it requires the fewest probe points per cycle (11 points for {1, 5, 10, . . . , 50} versus 17 for T s = 3), further reducing query overhead while maintaining near-optimal accuracy. 6.2. Multiple Spot Instances Dataset Analysis To address RQ-2, this section analyzes the spatial and temporal characteristics of the collected multi-node SPS dataset across both short-term and long-term horizons, followed by further analyses on instance size correlation, resource pool diversity, and T 3 characteristics. Short-term Spatial and Temporal Characteristics. The dataset covers July 1, 2024 to August 31, 2025, with 952 unique instance types across 17 AWS regions. Figure 6a and 6b show the average T 3 values during the week (August 4–11, 2025). A consistent daily cyclical pattern appears across all major regions, where T 3 values are higher during local nighttime and lower during
Correlation of T3 Values and Instance Size. To evaluate the correlation between the instance size and multi-node availability, T 3 scores of 5,414 pairs of adjacent-sized instances from the same family (e.g., m5.large and m5.2xlarge) were compared 8
41.0%
Smaller instance higher T3
40.1%
Small and large same T3
18.9%
Larger instance higher T3
Six Major Regions
30 20 10 0 0
(a) The distribution of correlation value for (b) Proportion of time with higher T3 values the T3 and the instance sizes across instance sizes
10
20 30 Difference
40
50
24-hour T3 sustatining ratio (%)
-0.5 0.0 0.5 1.0 Correlation Coefficient
100 80 60 40 20 0
Ratio (%)
CDF of Correlation Proportion (%)
CDF
1.0 0.8 0.6 0.4 0.2 0.0 -1.0
80 70 60 50 40 30 20 10 1 5 10 15 20 25 30 35 40 45 50 T3
Figure 9: T3 differences for the same Figure 10: The ratio of sustaining T3 instance type across regions and AZs values after 24 hours
SPS
Figure 7: Correlation of T3 values with respect to different instance sizes
3.0 2.5 2.0 1.5 1.0
Characteristics of T3. Figure 9 examines spatial variation of multi-node availability for each instance type by calculating the difference between its maximum and minimum T 3 values across all supported AZs whose value is shown in the horizontal axis. While some instance types show little variation, over 36% exhibit the maximum possible difference of 50, which means that they have at least one AZ with T 3 = 50 and another with T 3 ≈ 0. The choice of region and AZ is therefore a critical determinant to enhance multi-node stability. Figure 10 plots the proportion of instances that sustain their initial T 3 value over 24 hours, using 33,713 instance-type observations on August 1, 2025. The result follows a J-shaped curve, where for T 3 values between 1 and 45, higher initial values are less likely to be sustained (e.g., 70% sustaining ratio at T 3=1 versus only 15% at T 3=40), suggesting that moderately high availability is often transient. This trend reverses sharply at T 3=50, where the sustaining ratio reaches 74.1%. This anomaly is likely a ceiling effect of the 50-node query limit where instances whose true capacity far exceeds 50 are capped at T 3=50 and remain there even if their actual availability fluctuates above the threshold. While this is a limitation of the dataset, it also implies that a measured T 3 of 50 is a signal of high stability.
50 60 80 100 120 160 200 240 320 360 400 480 640 720 800 960 280 1 Total Core
Figure 8: SPS value distribution diversity to build a compute resource pool
using the August 2025 dataset. Figure 7a shows that 83.7% of pairs exhibit a positive T 3 correlation, indicating that availability patterns within the same family are generally synchronized across sizes; when the availability of smaller instances increases, the availability of larger instances tends to increase as well. However, synchronized patterns do not imply equal availability. Figure 7b shows that smaller instances had a higher T 3 value 41.0% of the time, while larger instances were superior only 18.9% of the time, with the remaining 40.1% being identical. Although smaller instances tend to offer better multi-node availability, larger instances still match or exceed them 59% of the time. Considering that using a larger instance type requires fewer nodes to meet a compute demand, instance size alone is not a reliable predictor, underscoring the need to reference per-instance-type T 3 rather than relying on size-based heuristics.
6.3. Effectiveness of Availability Scoring Mechanism This section addresses RQ-3 by validating the effectiveness of the proposed availability scoring mechanism. Experimental Setup. To measure real-world spot instance stability, experiments were conducted using 100 distinct instance types chosen to represent a wide range of predicted availability scores. From September 13 to October 9, 2024, 50 spot instances for each type were requested every 10 minutes, 24 hours a day. From the resulting spot instance request logs, a ground-truth metric, termed the Real Availability Score, was calculated following the methodology proposed by Wu et al. [45]. This real-world score was then compared against two predictors: 1. The proposed composite Predicted Availability Score. 2. A baseline using the raw, single-point vanilla T 3 value.
Diverse Spot Instance Availability. Rather than pre-selecting a specific instance type, a more effective strategy is to define a total resource requirement (e.g., target CPU cores) and explore the diverse instance combinations that can fulfill it. Figure 8 plots the SPS distribution for various instance combinations capable of fulfilling different total core count requirements. The median SPS shows an overall downward trend as the total number of required cores grows, falling below 2.0 around 120 cores and approaching 1.0 beyond 320 cores. Crucially, even for large requests, the upper quartiles and outliers reveal that high-SPS combinations still exist, though they might be harder to discover among the many low-availability options. This motivates the need for a recommendation engine that can identify these stable combinations from the multi-node SPS data.
Quantifying Effectiveness of the Availability Scoring Algorithm. Figure 11 compares the performance of these two predictors. The horizontal axis categorizes predicted availability score as Low (< 20), Mid (20–70), and High (> 70). The vertical axis shows the real availability score of instance types in each category based on the predicted scores. The proposed scoring heuristic 9
50 25 0
Low Mid High Predicted Availability Score
(a) Proposed availability score
100 75 50
Survival Rate
75
Real Availability Score
Real Availability Score
100
25 0
Low Mid High Predicted Availability Score
(b) Using vanilla SPS datasets (T 3) without score calculation
Figure 11: Superb real availability modeling capability of the proposed reliability scoring mechanism
75 Availability Score 50 Availability Score < 75 25 Availability Score < 50 Availability Score < 25
5
10 15 Time (hour)
20
25
Figure 12: The survival rate of spot instances with a different maximum number of high availability nodes
(Figure 11a) and the baseline (Figure 11b) show a positive correlation between their predicted scores and the real availability. However, a critical difference emerges in the Low score category where the baseline predictor incorrectly assigns many stable instances low scores and thus exhibits poor recall in finding stable spot instances; 26.3% of instances it categorized as Low were, in fact, highly available. In contrast, the proposed heuristic demonstrates superior recall, with a misclassification error rate of 11.1% in the same category. This result confirms that the proposed availability scoring mechanism, which incorporates temporal characteristics like trend and volatility, provides a significantly more accurate and reliable model of real-world spot instance stability than a predictor based on a single, unprocessed metric. After collecting the stability experiment data, we applied the Kaplan-Meier estimator (KME) [52] and the Cox proportional hazards model [53] to quantitatively evaluate spot instance survival times with the proposed availability score. These methods, commonly used to estimate survival rates or analyze variable impacts on survival [54, 55, 56], are widely applied in fields like medicine and business where the customer retention rate is important. Given their effectiveness in modeling lifetime data, they are well-suited for analyzing spot instance stability. The hazard ratio of the Cox proportional model is calculated as follows. h(t|x) = h0 (t) exp((x − x̄)′ β)
1.0 0.8 0.6 0.4 0.2 0.0 0
b S (t) =
Y ni − d i ! ni i: t ≤t
(6)
i
S (t) represents the probability that a The survival function b spot instance runs beyond time t. Here, ni is the number of running instances at time ti , and di is the number of interrupted instances. Figure 12 compares b S (t) across different availability score ranges. The x-axis represents runtime, while the y-axis shows the survival rate. Availability scores are grouped into four bins: solid lines indicate scores of 75 or higher, dasheddotted lines represent scores between 50 and 75, and so on. Higher scores correspond to longer survival times—instances with scores below 25 have a median survival time of 13 hours, while those scoring 75+ last 21.6 hours. This confirms that the proposed availability score effectively enhances spot instance reliability. Sensitivity of the Scaling Coefficient λ. The availability score (Eq. 3) includes a scaling coefficient λ that controls the magnitude of the trend and volatility adjustment. To validate the choice of λ = 0.1, a sensitivity analysis was conducted by sweeping λ from 0.0 to 1.0 in increments of 0.1. The accuracy was measured as the agreement between the predicted availability scores and the real availability scores derived from ground-truth interruption data, with improvements computed relative to the unadjusted baseline at λ = 0.0. Figure 13 plots the accuracy improvement as a function of λ. The accuracy improvement reaches its peak at λ = 0.1 with an improvement of +2.5 percentage points over the baseline. For λ ≥ 0.2, the adjustment over-amplifies estimation noise and the accuracy falls below the baseline across the remaining range. These results identify λ = 0.1 as the only operating point at which the trend and volatility adjustment improves the agreement between predicted and real availability scores, justifying its selection as the default coefficient in the proposed scoring model.
(5)
h(t|x) represents the hazard ratio at time t given the availability score x, while h0 (t) denotes the baseline hazard ratio without considering availability. β is the regression coefficient, estimated using the log-likelihood function [57], to quantify the impact of availability on spot interruptions. The results show a strong correlation between availability score and survival ratio (P ≤ 0.05), with a hazard ratio of 0.9903 (95% confidence interval: 0.9899–0.9907). Each 1-point increase in availability score reduces interruption risk by approximately 0.97%, following e−0.0097×∆x . At an availability score of 100, the risk decreases by about 62.1% compared to a score of 0, confirming its effectiveness as a reliability indicator. Next, to visualize the availability of multi-node spot instances based on the proposed availability score, we estimated their survival rate using the KME, calculated as follows:
Impact of T3 Observation Period. The length of the observation window used to compute the availability score from the T 3 timeseries directly affects score stability. To determine an appropriate window size, a sensitivity analysis was conducted by measuring 10
0
30 25 20 15 10 5 0 0.0
Ratio (%)
Accuracy Improvement over = 0 (%p)
5 5 10 15
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
0.2 0.4 0.6 0.8 Absolute Value of Correlation
1.0
Figure 15: High correlations of T 3 and T 2
| AS|
Spring
(2025/03~04)
Fall
(2025/09~10)
Summer
(2025/06~07)
Winter
(2025/12~2026/01)
100 90 80 70 60 50 40
Cost Score
8 6 4 2 0 8 6 4 2 0
100 80 60 40 20 0
Availability Score
| AS|
Figure 13: Sensitivity of the availability score to the scaling coefficient λ.
0.0
0.5 Weight
(a) Availability score
1.0
0.0
0.5 Weight
1.0
(b) Cost score
Figure 16: Impact of weight (W) to score distribution
2h 7h 1d 2d 4d 8d 14d 2h 7h 1d 2d 4d 8d 14d 1h 6h 12h 1d 3d 7d 3d 1h 6h 12h 1d 3d 7d 3d 1 1 Window Transition Window Transition
and T 2 are highly synchronized, and an availability score derived solely from T 3 is sufficient for assessing instance stability. Impact of Weight Parameter, W. To analyze the effect of W on ranking, 360 synthetic scenarios were constructed with varying vCPU (80–640) and memory (160–1280 GB) requirements across different instance categories, families, and types. Availability and cost scores were computed using data from August 17–22, 2025, with W set to 0.0, 0.5, and 1.0. Figure 16 shows the score distributions for the top-ranked instances under each setting. At W=0.0 (cost-only), all selected instances achieve a perfect cost score of 100, but their availability scores are low and widely dispersed. At W=1.0 (availabilityonly), availability scores are near-perfect, but cost scores drop substantially. At W=0.5 (balanced), the engine achieves availability scores nearly identical to the W=1.0 case while maintaining high cost-efficiency, demonstrating that a balanced weight provides near-optimal availability with only a minor cost compromise. This justifies W=0.5 as the default setting.
Figure 14: Distribution of |∆AS | across window transitions for four seasons.
|∆AS |, the absolute score difference between two consecutive window sizes when calculating the availability score, for the same base date. Figure 14 presents the |∆AS | distribution across seven window transitions over four seasonal periods from March 2025 to January 2026, each covering approximately 6,000 instancetransition observations. A consistent pattern emerges across all seasons: |∆AS | peaks at the 12h→1d transition, indicating that a full day of data is the critical threshold for capturing the daily cyclical patterns. Beyond one day, |∆AS | drops sharply, and by the 7d→8d transition, the median approaches near-zero. The consistency of this convergence behavior across all four seasons confirms that it is an inherent property of the T 3 data rather than a seasonal artifact. Based on this analysis, a seven-day observation window is adopted as the default for availability score computation. This duration captures weekly periodicity while remaining responsive to meaningful shifts in spot instance availability.
Effectiveness of Diversifying Instance Types. Relying on a single instance type for a large spot request increases interruption risk, motivating the proposed heuristic for constructing heterogeneous pools. Table 2 reports the [minimum, median, maximum] number of distinct instance types selected across scenarios with varying resource requirements (vCPU: 80–640; Memory: 160– 1280 GB) and three candidate pool scopes: broad (Category), restricted (Family), and narrow (Types). The results show that the algorithm adaptively adjusts pool diversity based on both the scale of the request and the breadth of available candidates. Figure 17 examines whether this diversification degrades recommendation quality. As instance types are progressively added to the pool, the average score declines only marginally in both broad (Figure 17a) and restricted (Figure 17b) candidate settings. This confirms that the heuristic’s termination conditions stop
Validity of Using Only T3. To verify that incorporating T 2 values when calculating the availability score is unnecessary, two availability scores were computed for each instance type over a one-week period (August 4–11, 2025): one from its T 3 timeseries and another from its T 2 time-series. Figure 15 shows the distribution of Pearson correlation coefficients [58] between the two scores. The distribution is heavily right-skewed, with approximately 25% of instances exhibiting near-perfect correlation (coefficient close to 1.0), while instances with low correlation (below 0.6) are infrequent. This confirms that the values of T 3 11
vCPU 80
160
Table 3: Execution time and score comparison between the greedy heuristic and ILP across candidate space scales
Memory
320
640
160
320
640
1280
Category [2,4,5] [2,3,5] [3,4,8] [4,4,7] [2,4,5] [2,5,8] [3,3,4] [4,4,7] Family [1,4,7] [2,3,8] [2,3,8] [3,5,7] [2,4,7] [2,4,8] [2,3,7] [3,4,8] Types [3,3,3] [3,4,4] [4,4,4] [4,4,4] [3,3,3] [3,4,4] [3,4,4] [4,4,4]
Time (ms)
Table 2: The number of minimum, median, and maximum types used in instance combinations for each scenario.
2
3
4
5
6
7
8
Number of Added Instance Types (a) Instance Category
9
100 80 60 40 20 0 1
Average Accumulated Score
Average Accumulated Score
100 80 60 40 20 0 1
2
3
4
5
6
7
8
Number of Added Instance Types
Sum of Score
Regions
Candidates
Greedy
ILP
Greedy
ILP
1 4 10 17
808 4,584 14,298 33,279
2.3 2.3 2.4 3.0
154 607 2,539 24,725
7,415 8,072 8,001 8,000
8,059 8,061 8,046 8,024
top-ranked candidates. The ILP solver exhibits rapidly accelerating execution time, which is from 154ms at 808 candidates to 24.7 seconds at 33,279 candidates, a 160× increase in runtime for a 41× increase in candidate count. At full scale, the ILP is over 8,000× slower than the greedy heuristic. In terms of solution quality, the score gap is at most 0.3% at full scale. At intermediate scales (4 and 10 regions), the greedy heuristic occasionally matches or slightly exceeds the ILP score; this occurs because the two methods optimize diversity differently, where the greedy heuristic uses score-proportional allocation while the ILP uses a linear diversity bonus, so neither strictly dominates the other in all cases. Overall, the greedy heuristic achieves comparable solution quality at a fraction of the computational cost, confirming its suitability for real-time recommendation over large candidate spaces.
9
(b) Instance Family
Figure 17: Average total scores as we diversify a set of recommended instance types
diversification before lower-ranked additions substantially degrade pool quality, achieving increased resilience with minimal score compromise. 6.3.1. Validation of Greedy-based Recommendation Approach To validate the effectiveness of the proposed greedy heuristic for recommendation, we compare it against an ILP formulation that jointly maximizes the total score and instance-type diversity.
6.4. Effectiveness of Instance Recommendation To address RQ-4, the effectiveness of SpotVista is evaluated against the state-of-the-art multi-region spot instance recommendation system, SpotVerse [27], and a public commercial service, AWS SpotFleet [60] in terms of cost and reliability.
ILP Formulation. To construct a reasonable baseline for comparison, we formulate an ILP that encodes both score maximization and type diversification as a single objective. While this formulation is not intended as a general-purpose optimal solver, it provides a useful reference point for assessing the quality gap of the greedy heuristic. For each candidate instance i, an integer variable xi denotes the number of allocated nodes, and a binary variable zi indicates whether instance type i is selected. The objective maximizes P P i S i · CPUi · xi + γ i zi , where the first term measures the vCPU-weighted pool quality and the second term rewards type diversity with coefficient γ. Linking constraints enforce zi = 1 whenever xi > 0, and a resource constraint ensures Rreq ≤ P i CPUi · xi ≤ Rreq + 1 to minimize over-provisioning. The binary variables expand the search space exponentially with the number of candidates. PuLP [59] with the CBC solver is used without a time limit.
Experimental Setup. SpotVerse recommends stable, costefficient spot instances in a multi-region setup, using SPS and Interruption-Free (IF) scores. The IF score represents the ratio of interruptions over the past 30 days, ranging from 1 to 3. SpotVerse sums SPS and IF scores, filtering instances with a total score above a threshold (default T = 4). It then selects the cheapest instance among candidates. To observe cases where availability is prioritized over cost, we also tested with T = 6. For a fair comparison, we conducted experiments in four major AWS regions (us-west-2, us-east-1, eu-west-2, ap-northeast-1) following SpotVerse’s setup; requesting the amount of compute and memory resources of 40 × m5.xlarge instances per region with thresholds T = 4 and T = 6. The interruption modeling experiment [45] ran for 24 hours from Feb. 3 to Feb. 5, 2025. While SpotVista normally recommends a mix of instance types to enhance stability, we constrained it to a single instance type per experiment to align with SpotVerse’s methodology that does not support diversifying instance types.
Setup and Results. Six days of SPS data (August 17–22, 2025) are used with a fixed requirement of 160 vCPU and W=0.5. The candidate set is scaled by progressively adding AWS regions from 1 to 17, growing the candidate instance types from 808 to 33,279. Table 3 reports execution time and total pool score at four representative scales. The user requirement is fixed at 160 vCPU, and the availability-cost weight W is set to 0.5. PuLP with the CBC solver is used for ILP execution. The greedy heuristic runs in approximately 2–3 ms regardless of scale, owing to the early termination after exploring only the
Comparing with SpotVerse. Figure 18 presents the comparative results for incurred hourly cost and instance availability. The analysis reveals that SpotVista consistently outperforms both SpotVerse configurations. In terms of cost (Figure 18a), SpotVista is the most efficient approach, reducing costs by 2.84% compared to SpotVerse-T4 12
2.5 2.0 1.5 1.0 0.5 0.0
SpotVerse - T6
Available Time (%)
Cost per Hour ($)
SpotVerse - T4
us-east-1
us-west-2 ap-northeast-1 eu-west-2 (a) Incurred cost per hour
100 75 50 25 0
SpotVista
us-east-1
us-west-2 ap-northeast-1 eu-west-2
(b) Spot instance available time percentage
86.0 82.5 75 61.7 58.3 57.0 59.7 53.7 53.7 50 25 0 SPS T3 LP PCO CO W:0.0 W:0.5 W:1.0 SingleTimepoint
SpotFleet
Availability (%)
Savings (%)
Figure 18: Comparing the spot instance recommendation quality of SpotVista with SpotVerse
SpotVista
75 62.3 70.4 50 44.8 51.4 48.8 49.6 49.6 47.1 25 0 SPS T3 LP PCO CO W:0.0 W:0.5 W:1.0 SingleTimepoint
(a) Cost savings ratio
SpotFleet
SpotVista
(b) Availability
Figure 19: Comparing SpotVista with AWS SpotFleet and simple approaches without using historical dataset
and 20.02% compared to the more expensive SpotVerse-T6. This is because SpotVista’s balanced scoring model can identify cost-efficient options that SpotVerse’s availability-first filtering mechanism overlooks. Figure 18b compares the percentage of time during which spot instances remained stable. SpotVerse-T6 places greater emphasis on spot instance reliability, resulting in significantly higher stability compared to SpotVerse-T4. The availability performance of SpotVista closely aligns with that of SpotVerse-T6. Overall, SpotVista achieves 1.53% higher availability than SpotVerse-T6 and 81.28% higher than SpotVerse-T4. In summary, SpotVista simultaneously achieves superior costefficiency and state-of-the-art availability, outperforming the existing state-of-the-art in both dimensions. Comparisons with older spot price prediction strategies [16, 17, 18, 19] are excluded, as recent studies have shown their reduced effectiveness due to changes in cloud vendor pricing policies [20, 21].
When multiple instances had the same highest values, we selected the lowest-priced one. For evaluation, we requested 50 requests every 10 minutes over a 24-hour period and recorded the success rate. Figure 19 presents a bar plot comparing availability and cost across different instance selection strategies. The horizontal axis represents the strategies, while the vertical axis shows cost savings (Figure 19a) and availability (Figure 19b). In Figure 19a, cost savings increase as SpotVista’s weight (W) decreases. SpotFleet’s lowest price (LP) strategy achieves slightly higher savings than its other strategies but is 30.6% lower than SpotVista at W = 0. The single time-point SPS approach also achieves high cost savings by selecting the cheapest instances with an SPS of three but still falls short of SpotVista’s cost efficiency. Figure 19b shows that SpotVista improves availability as W increases. SpotFleet’s CO and PCO strategies provide slightly better availability (0.8%) than LP but are 29.5% lower than SpotVista at W = 1. CO and PCO perform identically because SpotFleet makes the same recommendations, and their availability remains similar to LP, failing to meet expectations. All single time-point strategies show lower availability than SpotVista at W = 0.5 and W = 1.0.
Comparing with AWS SpotFleet. Next, we compare SpotVista with AWS SpotFleet [60], which enables users to configure a pool of spot instances using a launch template and adjust allocation strategies for cost and stability. SpotFleet supports three strategies: Lowest Price (LP), Capacity Optimized (CO), and Price-Capacity Optimized (PCO). SpotVista can emulate these strategies by adjusting W: W = 0.0 for LP, W = 1.0 for CO, and W = 0.5 for PCO. We evaluate all configurations. Due to SpotFleet’s regional constraints, experiments were conducted in us-east-1, one of the largest AWS regions. Additionally, we compared performance against a naive approach that selects multi-node spot instances based solely on single-node SPS and T 3 values at the request time, ignoring temporal effects.
Overall, compared to AWS SpotFleet, SpotVista improves availability by over 20% while maintaining similar cost savings. Additionally, when availability is comparable, SpotVista achieves over 25% more cost savings, demonstrating its effectiveness and practicality. 13
7. Related Work
seasonality and higher amplitude variability than AWS. In addition, the data collected from the Azure API contains missing and inconsistent responses, which hinders continuous time-series collection. A complete extension would require additional preprocessing such as missing-value imputation, together with an adapted scoring model that accounts for weaker periodic structure. GCP and Alibaba do not currently expose a public availability score API, but the proposed methodology can be extended to these platforms once comparable interfaces become available.
Modeling and Utilization of Spot Instance Datasets: Since the introduction of spot instances, many attempts have been made to correlate spot instance price datasets with interruption risks to minimize the likelihood of instance termination [61, 2, 62, 3, 4, 5, 63]. For example, Ali-Eldin et al. proposed a heuristic of the deployment of web servers based on spot prices and analysis [9], while Lee et al. introduced DeepSpotCloud [6], which uses GPU spot instances across global regions for DNN training tasks. In big data processing, SeeSpotRun [7] suggested methods for efficiently using spot instances in a Hadoop [64] cluster, while Flint [8] and Tr-Spark [65] were proposed for Apache Spark [66]. Fabra et al. [19] proposed a DNN model to predict spot instance prices to enhance its reliability. However, these previous works relied on historical spot price datasets, which became obsolete after the operational changes [20, 21]. This paper presents a new approach that does not depend on spot price datasets, enabling the selection of cost-efficient and stable multiple spot instances in a different way. Spot Instance Interruption Analysis: Pham et al.[44], Lee et al.[24], and Kim et al.[36] conducted experiments on AWS spot instances, analyzing interruption patterns. In the case of Azure, Yang et al.[67] proposed an interruption prediction model based on the Transformer architecture, using Azure’s internal interruption tracking data. For GCP, Haugerud et al.[68] and Kadupitiya et al.[69] modeled spot instance interruptions. Additionally, Yang et al. [70] performed interruption experiments in multi-cloud environments. Our proposed approach is complementary to previous research, as it leverages a new type of spot instance dataset to improve reliability. Utilization of Multi-Node Spot Instances: As the scale for computing resources increases for various applications, such as deep learning and big data processing, several approaches have been proposed to reliably utilize multiple spot instance nodes, focusing on reducing training costs while handling interruptions and improving training throughput [33, 34, 71, 72, 73]. We expect that the algorithm proposed in this work will increase the utility of these previous research outcomes. Yang et al.[49] and Xu et al.[74] proposed methods to build cost-efficient and reliable clusters by mixing on-demand and spot instances. Sharma et al.[75] and Harlap et al.[76] introduced frameworks that balance cost and reliability for spot instance clusters. While existing work relies on internal datasets or spot price datasets, to the best of our knowledge, this paper is the first to quantitatively model the availability of large-scale spot instances using public datasets other than price.
Reactive Adjustment after Deployment. SpotVista currently operates as a one-shot recommendation engine and does not monitor spot instance availability or cost score changes after deployment. A natural extension is to integrate SpotVista as an availability signal provider for workload-agnostic cluster managers such as SkyPilot [70], enabling continuous rebalancing throughout the workload lifetime. Such an extension would require a reactive decision loop that periodically reevaluates the active pool against updated availability signals, which we consider a promising direction for future research. 9. Conclusion This paper addressed the critical challenge of provisioning reliable multi-node workloads on volatile cloud spot instances, a problem for which single-node availability metrics are inadequate. To solve this challenge, we proposed SpotVista whose key contributions include an efficient, sampling-based heuristic for collecting large-scale multi-node availability data, a new scoring model that quantifies instance stability by analyzing temporal patterns, and a recommendation engine that constructs diverse, cost-efficient, and reliable instance pools. Through extensive real-world experiments, SpotVista demonstrated significant performance improvement over existing solutions. Compared to the state-of-the-art research system, SpotVerse [27], it achieved 81.28% greater reliability and 2.84% higher cost-efficiency. Furthermore, it provided 21.6% higher stability than the commercially available AWS SpotFleet service, underscoring its practical effectiveness. To benefit the research community and cloud practitioners, the collected multi-node availability dataset and the recommendation engine have been made publicly available through a web service. Future work will focus on extending the system to support dynamic, post-recommendation adjustments and expanding compatibility to other major cloud vendors. Acknowledgements This work was supported by Institute of Information & communications Technology Planning & Evaluation(IITP) grant funded by the Korea government(MSIT) (RS-2022-00144309 & RS-2025-25441560 & RS-2026-25492200)
8. Discussion Generalization to Other Cloud Vendors. The methodology of SpotVista is vendor-agnostic, but its applicability requires two prerequisites. A queryable availability indicator analogous to AWS SPS must be exposed, and the collected data must support reliable historical pattern-based scoring. Azure has recently begun offering an SPS API [23], and the MSTL analysis in Section 6.2 shows that Azure T 3 exhibits substantially weaker
References [1] O. Agmon Ben-Yehuda, M. Ben-Yehuda, A. Schuster, D. Tsafrir, Deconstructing amazon ec2 spot instance pricing, ACM Trans. Econ. Comput. 1 (3) (sep 2013). doi:10.1145/2509413.2509416.
14
[2] D. Movsowitz Davidow, O. Agmon Ben-Yehuda, O. Dunkelman, Deconstructing alibaba cloud’s preemptible instance pricing, in: Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing, HPDC ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 253–265. doi:10.1145/3588195.3593 001. [3] B. Javadi, R. K. Thulasiram, R. Buyya, Characterizing spot price dynamics in public cloud environments, Future Generation Computer Systems 29 (4) (2013) 988–999, special Section: Utility and Cloud Computing. doi: https://doi.org/10.1016/j.future.2012.06.012. [4] C. Wang, Q. Liang, B. Urgaonkar, An empirical analysis of amazon ec2 spot instance features affecting cost-effective resource procurement, in: Proceedings of the 8th ACM/SPEC on International Conference on Performance Engineering, ICPE ’17, Association for Computing Machinery, New York, NY, USA, 2017, p. 63–74. doi:10.1145/3030207.3030210. [5] N. Ekwe-Ekwe, A. Barker, Location, location, location: Exploring amazon ec2 spot instance pricing across geographical regions, in: 2018 18th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), 2018, pp. 370–373. doi:10.1109/CCGRID.2018.0005 9. [6] K. Lee, M. Son, Deepspotcloud: Leveraging cross-region gpu spot instances for deep learning, in: 2017 IEEE 10th International Conference on Cloud Computing (CLOUD), 2017, pp. 98–105. doi:10.1109/CLOUD. 2017.21. [7] N. Chohan, C. Castillo, M. Spreitzer, M. Steinder, A. Tantawi, C. Krintz, See spot run: Using spot instances for MapReduce workflows, in: 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 10), USENIX Association, Boston, MA, 2010. URL https://www.usenix.org/conference/hotcloud-10/see-spot-run-usi ng-spot-instances-mapreduce-workflows [8] P. Sharma, T. Guo, X. He, D. Irwin, P. Shenoy, Flint: Batch-interactive dataintensive processing on transient servers, in: Proceedings of the Eleventh European Conference on Computer Systems, EuroSys ’16, Association for Computing Machinery, New York, NY, USA, 2016. doi:10.1145/2901 318.2901319. [9] A. Ali-Eldin, J. Westin, B. Wang, P. Sharma, P. Shenoy, Spotweb: Running latency-sensitive distributed web services on transient cloud servers, in: Proceedings of the 28th International Symposium on HighPerformance Parallel and Distributed Computing, HPDC ’19, Association for Computing Machinery, New York, NY, USA, 2019, p. 1–12. doi:10.1145/3307681.3325397. [10] X. He, P. Shenoy, R. Sitaraman, D. Irwin, Cutting the cost of hosting online services using cloud spot markets, in: Proceedings of the 24th International Symposium on High-Performance Parallel and Distributed Computing, HPDC ’15, Association for Computing Machinery, New York, NY, USA, 2015, p. 207–218. doi:10.1145/2749246.2749275. [11] I. Menache, O. Shamir, N. Jain, On-demand, spot, or both: Dynamic resource allocation for executing batch jobs in the cloud, in: 11th International Conference on Autonomic Computing (ICAC 14), USENIX Association, Philadelphia, PA, 2014, pp. 177–187. URL https://www.usenix.org/conference/icac14/technical-sessions/pres entation/menache [12] S. Subramanya, T. Guo, P. Sharma, D. Irwin, P. Shenoy, Spoton: A batch computing service for the spot market, in: Proceedings of the Sixth ACM Symposium on Cloud Computing, SoCC ’15, Association for Computing Machinery, New York, NY, USA, 2015, p. 329–341. doi:10.1145/28 06777.2806851. [13] P. Varshney, Y. Simmhan, Autobot: Resilient and cost-effective scheduling of a bag of tasks on spot vms, IEEE Transactions on Parallel & Distributed Systems 30 (07) (2019) 1512–1527. doi:10.1109/TPDS.2018.2889 851. [14] L. Zheng, C. Joe-Wong, C. W. Tan, M. Chiang, X. Wang, How to bid the cloud, in: Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication, SIGCOMM ’15, Association for Computing Machinery, New York, NY, USA, 2015, p. 71–84. doi:10.1145/2785956.2787473. [15] P. Sharma, D. Irwin, P. Shenoy, How not to bid the cloud, in: 8th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 16), USENIX Association, Denver, CO, 2016. URL https://www.usenix.org/conference/hotcloud16/workshop-program /presentation/sharma
[16] S. Alkharif, K. Lee, H. Kim, Time-series analysis for price prediction of opportunistic cloud computing resources, in: W. Lee, W. Choi, S. Jung, M. Song (Eds.), Proceedings of the 7th International Conference on Emerging Databases, Springer Singapore, Singapore, 2018, pp. 221–229. [17] V. Khandelwal, A. Chaturvedi, C. P. Gupta, Amazon ec2 spot price prediction using regression random forests, IEEE Transactions on Cloud Computing (2017) 1–1doi:10.1109/TCC.2017.2780159. [18] M. Khodak, L. Zheng, A. S. Lan, C. Joe-Wong, M. Chiang, Learning cloud dynamics to optimize spot instance bidding strategies, in: IEEE INFOCOM 2018 - IEEE Conference on Computer Communications, 2018, pp. 2762–2770. doi:10.1109/INFOCOM.2018.8486291. [19] J. Fabra, J. Ezpeleta, P. Álvarez, Reducing the price of resource provisioning using ec2 spot instances with prediction models, Future Generation Computer Systems 96 (2019) 348–367. doi:https://doi.org/10.1 016/j.future.2019.01.025. [20] M. Baughman, S. Caton, C. Haas, R. Chard, R. Wolski, I. Foster, K. Chard, Deconstructing the 2017 changes to aws spot market pricing, in: Proceedings of the 10th Workshop on Scientific Cloud Computing, ScienceCloud ’19, Association for Computing Machinery, New York, NY, USA, 2019, p. 19–26. doi:10.1145/3322795.3331465. [21] D. Irwin, P. Shenoy, P. Ambati, P. Sharma, S. Shastri, A. Ali-Eldin, The price is (not) right: Reflections on pricing for transient cloud servers, in: 2019 28th International Conference on Computer Communication and Networks (ICCCN), 2019, pp. 1–9. doi:10.1109/ICCCN.2019.88469 33. [22] A. W. is New, Introducing amazon ec2 spot placement score (2021). URL https://aws.amazon.com/about-aws/whats-new/2021/10/amazon-e c2-spot-placement-score/ [23] Azure, Spot placement score (2025). URL https://learn.microsoft.com/en-us/azure/virtual-machine-scale-set s/spot-placement-score [24] S. Lee, J. Hwang, K. Lee, Spotlake: Diverse spot instance dataset archive service, in: 2022 IEEE International Symposium on Workload Characterization (IISWC), IEEE Computer Society, Los Alamitos, CA, USA, 2022, pp. 242–255. doi:10.1109/IISWC55918.2022.00029. [25] K. Kim, S. Park, J. Hwang, H. Lee, S. Kang, K. Lee, Public spot instance dataset archive service, in: Companion Proceedings of the ACM Web Conference 2023, WWW ’23 Companion, Association for Computing Machinery, New York, NY, USA, 2023, p. 69–72. doi:10.1145/3543 873.3587314. [26] S. Cheon, K. Kim, K. Kim, M. Song, K. Lee, Multi-node spot instances availability score collection system, in: Proceedings of the 34th International Symposium on High-Performance Parallel and Distributed Computing (HPDC ’25), ACM, 2025. [27] M. Son, G. G. Akbulut, M. T. Kandemir, Spotverse: Optimizing bioinformatics workflows with multi-region spot instances in galaxy and beyond, in: Proceedings of the 25th International Middleware Conference, Middleware ’24, Association for Computing Machinery, New York, NY, USA, 2024, p. 74–87. doi:10.1145/3652892.3700750. [28] S. Tang, J. Yuan, X.-Y. Li, Towards optimal bidding strategy for amazon ec2 cloud spot instance, in: 2012 IEEE Fifth International Conference on Cloud Computing, 2012, pp. 91–98. doi:10.1109/CLOUD.2012.134. [29] azure cloud, azure spot virtual machines (2024). URL https://learn.microsoft.com/en-us/azure/virtual-machines/spot-vms/ [30] alibaba cloud, alibaba preemptible instance (2024). URL https://www.alibabacloud.com/help/en/ecs/user-guide/overview-4/ [31] R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley, Y. He, Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’22, IEEE Press, 2022. [32] J. Sampé, M. Sánchez-Artigas, G. Vernik, I. Yehekzel, P. Garcı́a-López, Outsourcing data processing jobs with lithops, IEEE Transactions on Cloud Computing 11 (1) (2023) 1026–1037. doi:10.1109/TCC.2021.31290 00. [33] J. Thorpe, P. Zhao, J. Eyolfson, Y. Qiao, Z. Jia, M. Zhang, R. Netravali, G. H. Xu, Bamboo: Making preemptible instances resilient for affordable training of large DNNs, in: 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), USENIX Association, Boston, MA, 2023, pp. 497–513. URL https://www.usenix.org/conference/nsdi23/presentation/thorpe
15
[34] J. Duan, Z. Song, X. Miao, X. Xi, D. Lin, H. Xu, M. Zhang, Z. Jia, Parcae: Proactive, Liveput-Optimized DNN training on preemptible instances, in: 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), USENIX Association, Santa Clara, CA, 2024, pp. 1121–1139. URL https://www.usenix.org/conference/nsdi24/presentation/duan [35] N. Chohan, C. Castillo, M. Spreitzer, M. Steinder, A. Tantawi, C. Krintz, See spot run: Using spot instances for {MapReduce} workflows, in: 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 10), 2010. [36] K. Kim, K. Lee, Making cloud spot instance interruption events visible, in: Proceedings of the ACM on Web Conference 2024, WWW ’24, Association for Computing Machinery, New York, NY, USA, 2024, p. 2998–3009. doi:10.1145/3589334.3645548. [37] AWS, Best practices for amazon ec2 spot (2026). URL https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-bes t-practices [38] R. M. Gray, Entropy and Information Theory, 2nd Edition, Springer Publishing Company, Incorporated, 2011. [39] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, E. Duchesnay, Scikit-learn: Machine learning in Python - minmax scaling, Journal of Machine Learning Research 12 (2011) 2825–2830. [40] A. Karalič, I. Bratko, First order regression, Machine learning 26 (1997) 147–176. [41] R. M. Karp, Reducibility among combinatorial problems, in: R. E. Miller, J. W. Thatcher (Eds.), Complexity of computer computations, Plenum Press, New York, NY, USA, 1972, pp. 85–103. [42] J. M. Hellerstein, J. M. Faleiro, J. Gonzalez, J. Schleier-Smith, V. Sreekanti, A. Tumanov, C. Wu, Serverless computing: One step forward, two steps back, in: 9th Biennial Conference on Innovative Data Systems Research, CIDR 2019, Asilomar, CA, USA, January 13-16, 2019, Online Proceedings, www.cidrdb.org, 2019. URL http://cidrdb.org/cidr2019/papers/p119-hellerstein-cidr19.pdf [43] J. Kadupitige, V. Jadhao, P. Sharma, Modeling the temporally constrained preemptions of transient cloud vms, in: Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing, HPDC ’20, Association for Computing Machinery, New York, NY, USA, 2020, p. 41–52. doi:10.1145/3369583.3392671. [44] T.-P. Pham, S. Ristov, T. Fahringer, Performance and behavior characterization of amazon ec2 spot instances, in: 2018 IEEE 11th International Conference on Cloud Computing (CLOUD), 2018, pp. 73–81. doi:10.1109/CLOUD.2018.00017. [45] Z. Wu, W.-L. Chiang, Z. Mao, Z. Yang, E. Friedman, S. Shenker, I. Stoica, Can’t be late: Optimizing spot instance savings under deadlines, in: 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), USENIX Association, Santa Clara, CA, 2024, pp. 185–203. URL https://www.usenix.org/conference/nsdi24/presentation/wu-zhang hao [46] Z. Li, T. Xia, Z. Mao, Z. Zhou, E. J. Jackson, J. Kerney, Z. Wu, P. Mishra, Y. Xu, Y. Qiao, S. Shenker, I. Stoica, Skynomad: On using multi-region spot instances to minimize ai batch job cost (2026). arXiv:2601.06520. URL https://arxiv.org/abs/2601.06520 [47] K. Kim, M. Song, T. Kim, K. Lee, Spot-and-scoot: Peeking into spot instance availability (Apr. 2026). doi:10.5281/zenodo.19590680. URL https://doi.org/10.5281/zenodo.19590680 [48] K. Bandara, R. J. Hyndman, C. Bergmeir, Mstl: A seasonal-trend decomposition algorithm for time series with multiple seasonal patterns, International Journal of Operational Research 52 (1) (2025) 79–98. [49] F. Yang, L. Wang, Z. Xu, J. Zhang, L. Li, B. Qiao, C. Couturier, C. Bansal, S. Ram, S. Qin, Z. Ma, I. n. Goiri, E. Cortez, T. Yang, V. Rühle, S. Rajmohan, Q. Lin, D. Zhang, Snape: Reliable and low-cost computing with mixture of spot and on-demand vms, in: Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, ASPLOS 2023, Association for Computing Machinery, New York, NY, USA, 2023, p. 631–643. doi:10.1145/3582016.3582028. [50] X. Wang, K. Smith, R. Hyndman, Characteristic-based clustering for time series data, Data Min. Knowl. Discov. 13 (3) (2006) 335–364. doi: 10.1007/s10618-005-0039-x.
[51] J. Bai, P. Perron, Estimating and testing linear models with multiple structural changes, Econometrica (1998) 47–78. [52] E. L. Kaplan, P. Meier, Nonparametric estimation from incomplete observations, Journal of the American Statistical Association 53 (282) (1958) 457–481. doi:10.1080/01621459.1958.10501452. [53] D. R. Cox, Regression models and life-tables, Journal of the Royal Statistical Society: Series B (Methodological) 34 (2) (1972) 187–202. [54] D. A. Ali, A. M. Hussein, Analysis of cox proportional hazard model for dropout students in university: case study from simad university, Journal of Applied Research in Higher Education 16 (3) (2024) 820–830. [55] K. N. Chi, T. Kheoh, C. J. Ryan, A. Molina, J. Bellmunt, N. J. Vogelzang, D. E. Rathkopf, K. Fizazi, P. W. Kantoff, J. Li, et al., A prognostic index model for predicting overall survival in patients with metastatic castrationresistant prostate cancer treated with abiraterone acetate after docetaxel, Annals of Oncology 27 (3) (2016) 454–460. [56] M. Okuda-Arai, S. Mori, F. Takano, K. Ueda, M. Sakamoto, Y. YamadaNakanishi, M. Nakamura, Impact of glaucoma medications on subsequent schlemm’s canal surgery outcome: Cox proportional hazard model and propensity score-matched analysis, Acta Ophthalmologica 102 (2) (2024) e178–e184. [57] D. Conniffe, Expected maximum log likelihood estimation, Journal of the Royal Statistical Society. Series D (The Statistician) 36 (4) (1987) 317–329. URL http://www.jstor.org/stable/2348828 [58] J. Benesty, J. Chen, Y. Huang, I. Cohen, Pearson correlation coefficient, in: Noise reduction in speech processing, Springer, 2009, pp. 1–4. [59] S. A. Mitchell, Pulp, https://github.com/coin-or/pulp (2003). [60] AWS, Ec2 fleet and spot fleet (2024). URL https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Fleets [61] R. Wolski, J. Brevik, R. Chard, K. Chard, Probabilistic guarantees of execution duration for amazon spot instances, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’17, Association for Computing Machinery, New York, NY, USA, 2017. doi:10.1145/3126908.3126953. [62] G. Portella, G. N. Rodrigues, E. Nakano, A. C. Melo, Statistical analysis of amazon ec2 cloud pricing models, Concurrency and Computation: Practice and Experience 31 (18) (2019) e4451. doi:10.1002/cpe.4451. [63] A. Marathe, R. Harris, D. Lowenthal, B. R. de Supinski, B. Rountree, M. Schulz, Exploiting redundancy for cost-effective, time-constrained execution of hpc applications on amazon ec2, in: Proceedings of the 23rd International Symposium on High-Performance Parallel and Distributed Computing, HPDC ’14, Association for Computing Machinery, New York, NY, USA, 2014, p. 279–290. doi:10.1145/2600212.2600226. [64] A. S. Foundation, Apache hadoop (2004). URL http://hadoop.apache.org/ [65] Y. Yan, Y. Gao, Y. Chen, Z. Guo, B. Chen, T. Moscibroda, Tr-spark: Transient computing for big data analytics, in: Proceedings of the Seventh ACM Symposium on Cloud Computing, SoCC ’16, Association for Computing Machinery, New York, NY, USA, 2016, p. 484–496. doi:10.1145/2987550.2987576. [66] M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauly, M. J. Franklin, S. Shenker, I. Stoica, Resilient distributed datasets: A FaultTolerant abstraction for In-Memory cluster computing, in: 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12), USENIX Association, San Jose, CA, 2012, pp. 15–28. URL https://www.usenix.org/conference/nsdi12/technical-sessions/pres entation/zaharia [67] F. Yang, B. Pang, J. Zhang, B. Qiao, L. Wang, C. Couturier, C. Bansal, S. Ram, S. Qin, Z. Ma, I. n. Goiri, E. Cortez, S. Baladhandayutham, V. Rühle, S. Rajmohan, Q. Lin, D. Zhang, Spot virtual machine eviction prediction in microsoft cloud, in: Companion Proceedings of the Web Conference 2022, WWW ’22, Association for Computing Machinery, New York, NY, USA, 2022, p. 152–156. doi:10.1145/3487553.3524229. [68] H. Haugerud, J. Krüger Svensson, A. Yazidi, Autonomous provisioning of preemptive instances in google cloud for maximum performance per dollar, in: 2020 5th International Conference on Cloud Computing and Artificial Intelligence: Technologies and Applications (CloudTech), 2020, pp. 1–8. doi:10.1109/CloudTech49835.2020.9365879. [69] J. Kadupitiya, V. Jadhao, P. Sharma, Scispot: Scientific computing on temporally constrained cloud preemptible vms, IEEE Transactions on Parallel and Distributed Systems 33 (12) (2022) 3575–3588. doi:10.1
16
109/TPDS.2022.3157272. [70] Z. Yang, Z. Wu, M. Luo, W.-L. Chiang, R. Bhardwaj, W. Kwon, S. Zhuang, F. S. Luan, G. Mittal, S. Shenker, I. Stoica, SkyPilot: An intercloud broker for sky computing, in: 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), USENIX Association, Boston, MA, 2023, pp. 437–455. URL https://www.usenix.org/conference/nsdi23/presentation/yang-zon gheng [71] I. Jang, Z. Yang, Z. Zhang, X. Jin, M. Chowdhury, Oobleck: Resilient distributed training of large models using pipeline templates, in: Proceedings of the 29th Symposium on Operating Systems Principles, SOSP ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 382–395. doi:10.1145/3600006.3613152. [72] S. Athlur, N. Saran, M. Sivathanu, R. Ramjee, N. Kwatra, Varuna: scalable, low-cost training of massive deep learning models, in: Proceedings of the Seventeenth European Conference on Computer Systems, EuroSys ’22, Association for Computing Machinery, New York, NY, USA, 2022, p. 472–487. doi:10.1145/3492321.3519584. [73] Y. Kim, K. Kim, Y. Cho, J. Kim, A. Khan, K.-D. Kang, B.-S. An, M.H. Cha, H.-Y. Kim, Y. Kim, Deepvm: Integrating spot and on-demand vms for cost-efficient deep learning clusters in the cloud, in: 2024 IEEE 24th International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2024, pp. 227–235. doi:10.1109/CCGrid59990.2024.000 34. [74] Z. Xu, C. Stewart, N. Deng, X. Wang, Blending on-demand and spot instances to lower costs for in-memory storage, in: IEEE INFOCOM 2016 - The 35th Annual IEEE International Conference on Computer Communications, 2016, pp. 1–9. doi:10.1109/INFOCOM.2016.75243 48. [75] P. Sharma, D. Irwin, P. Shenoy, Portfolio-driven resource management for transient cloud servers, Proceedings of the ACM on Measurement and Analysis of Computing Systems 1 (1) (2017) 1–23. [76] A. Harlap, A. Chung, A. Tumanov, G. R. Ganger, P. B. Gibbons, Tributary: spot-dancing for elastic services with latency SLOs, in: 2018 USENIX Annual Technical Conference (USENIX ATC 18), USENIX Association, Boston, MA, 2018, pp. 1–14. URL https://www.usenix.org/conference/atc18/presentation/harlap
17