1
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
Abstract—Cloud performance fluctuates due to factors such as resource contention and workload changes. These factors can be short-term, seasonal, or long-term. Their effects are often intertwined in performance traces, making performance management difficult. Prior work on cloud performance engineering used time-series decomposition to separate these factors. However, existing approaches rely on basic decomposition methods that may miss key variation patterns and fail on traces with complex or intermittent patterns, limiting their usefulness across diverse cloud deployments. To address this limitation, we propose two time-series decomposition techniques for cloud performance engineering: a hybrid/manual method and a fully automatic method. Through a case study of 11 serverless functions, we show that both approaches can successfully and consistently reveal trends and seasonal cycles, such as weekly and quarterly patterns, which are otherwise obscured. As an evaluation and application of the decomposition, we used the decomposed components to predict future performance, yielding mean absolute percentage error (MAPE) values of only 1.8% (hybrid) and 2.1% (automatic), significantly outperforming basic time-series methods and deep learning. We further show that decomposition insights can guide practical resource allocation. Using decomposition-informed scaling on AWS, we reduced latency variability by over 60% and maximum latency by 10%. Similar experiments on benchmarks on AWS confirmed that seasonal patterns and performance gains generalize beyond our case study. Notably, our findings demonstrate that even a single performance trace contains rich actionable information for guiding cloud management decisions. Index Terms—cloud performance, time-series decomposition, latency analysis, performance engineering
I. I NTRODUCTION Cloud-native applications (CNAs) are highly scalable applications designed for dynamic cloud environments [1], [2]. Their high scalability has made them the preferred choice for hosting large-scale web services [3], [4] and machine learning models [5], [6], [7], which often demand millisecondlevel latencies [8]. Thus, performance engineering—including debugging, optimization, monitoring, and analysis—is critical for CNAs both in development and production [9], [10], [11]. However, managing cloud performance is notoriously challenging due to significant performance fluctuations [12], [13], [14], [15]. For example, Fig. 1 illustrates the latency trace of CaptureStripe, a serverless CNA from the Serverless Airline Shimul Debnath, Donald Lien, and Wei Wang are with the University of Texas at San Antonio, USA. E-mail: [email protected], [email protected], [email protected]. William Hart and Lori Pollock are with the University of Delaware, USA. E-mail: [email protected], [email protected]. This work has been submitted to IEEE for possible publication.
Latency (ms)
arXiv:2605.09787v1 [cs.DC] 10 May 2026
Shimul Debnath, William Hart, Lori Pollock, Donald Lien, and Wei Wang
550 CaptureStripe 500 450 07/25/ 09/13/ 11/03/ 12/24/ 02/12/ 04/03/ 05/23/ 20 20 20 20 21 21 21
Fig. 1: Latency trace of a cloud function over 301 days (07/2020∼05/2021) on Amazon AWS Cloud.
Booking Application (SAB), recorded over 43 weeks (301 days) under a constant workload and resource configuration [16]. Despite its stable setup, the application’s latency increased substantially from 450ms to 550ms over the year. Such a drastic performance shift necessitates investigation and mitigation, both of which require identifying the cause of the latency increase. Since neither workload nor resource allocation changed, this degradation was likely driven by other factors, such as hardware or software updates in the data center or resource contention due to multi-tenancy. However, identifying root causes from performance traces without clear patterns (like the trace in Fig. 1) is extremely difficult. A major challenge is that cloud performance fluctuations are typically driven by multiple factors acting simultaneously. Their effects intertwine within a single trace, making it difficult to isolate individual effect for reliable analysis. Thus, a key requirement in cloud performance engineering is the ability to decompose a trace into distinct components, each representing a coherent variation pattern. Prior work has applied time-series decomposition to break cloud performance traces into components such as long-term trends, seasonal (temporal) cycles (e.g., quarterly or monthly), and other variations. For example, isolating long-term trends is widely used in the IT industry to detect performance regressions caused by code updates. [17], [18], [19]. While time-series decomposition is effective for detecting cloud performance regressions, simple techniques used in prior work, such as STL (Seasonal-Trend Decomposition using LOESS) [20], can not always accurately or fully identify the components within cloud performance traces [21], especially from public clouds, due to the following challenges. 1. Simple decomposition techniques, such as STL, report all cyclic variations as a single component, without distinguishing between weekly, monthly, or quarterly cycles. While this may suffice for performance regression detection, where only the trend is needed, it is often inadequate for other performance and resource management tasks.
2
2. Simple decompositions rely on predefined models with fixed parameters. For example, STL applies only LOESS regression with a fixed seasonal period (e.g., weekly). Such fixed-model approaches are less effective for public cloud traces, which often exhibit diverse characteristics (some even non-parametric) and varying periodicities. 3. Cloud performance variations are often non-stationary and intermittent due to the random multi-tenancy. For instance, a weekly cycle (e.g., high latency on Wednesdays) may vanish during holidays. Such irregularity, known as signal intermittency [22], [23], causes mode mixing in simple decomposition methods, where one component includes multiple variation patterns [24], [25]. To address these challenges, this paper explores advanced time-series analysis techniques for decomposing cloud performance traces. Specifically, we propose two decomposition methods: one hybrid/manual and one fully automatic. Our hybrid/manual decomposition follows a top-down approach, where users iteratively identify components, starting with the long-term trend and then seasonal patterns. The process continues til the residual appears random, ensuring all significant components extracted. Different model types, parametric or non-parametric, can be applied as needed (thus hybrid). Users may also fine-tune model parameters to reduce the impact of signal intermittency on automatic fitting. For users without statistical expertise or want fully automation, we provide an automatic decomposition technique. After evaluating several methods, we found Ensemble Empirical Mode Decomposition (EEMD) [26], [27] to be effective for cloud performance traces. EEMD iteratively extracts components until residuals have low variation, ensuring all relevant components are identified. As a data-driven method, EEMD is not limited by predefined models and is specifically designed to mitigate mode mixing from signal intermittency [27]. As a case study, we apply our decomposition techniques to performance traces from the Serverless Airline Booking Application (SAB) [16], which includes 11 serverless functions in Python, JavaScript, and TypeScript. Amazon developed SAB to showcase CNA designs using cloud-native services. This case study confirms the effectiveness of our decomposition techniques, as they consistently extract components corresponding to quarterly, monthly, and weekly cycles in performance traces. These findings suggest that cloud performance exhibits a degree of regularity and predictability, enabling performance and resource management strategies to be aligned with these periodic patterns. The results also show that our hybrid/manual and automatic decomposition techniques are largely equivalent, identifying nearly the same components despite their distinct mathematical foundations. As an application of our decomposition techniques, we developed performance prediction models for the SAB applications to forecast their performance over the next 28 days. These predictions achieved high accuracy, with average MAPE values of only 1.8% (hybrid/manual) and 2.1% (automatic). Besides low error rates, our models also effectively captured latency peaks and valleys across most applications. Moreover, our predictions are more accurate than the commonly used STL decomposition and a neural network model. These highly
accurate predictions also validate the correctness of our decomposition methods. As a second application, we applied our decomposition results to optimize SAB’s resource allocation for better performance and stability. Experiments results showed that decomposition-informed allocations reduced latency standard deviation by 60.2% and slowest latency by 10.8%, demonstrating that accurate performance decomposition can yield real benefits for cloud deployments. Additional experiments with other benchmarks on AWS and Google Cloud confirm that seasonal patterns and performance benefits generalize across platforms. To the best of our knowledge, this work is the first to demonstrate how to reliably and completely extract trends and temporal cycles from noisy and irregular public cloud traces. Prior to our study, it was unclear whether meaningful trends and temporal cycles could be fully and correctly extracted, and which decomposition techniques were appropriate for such extraction tasks. Our work is novel as we address these open questions by systematically evaluating decomposition methods and demonstrating how to robustly uncover meaningful patterns, enabling informed and practical cloud performance management. The key contributions of this paper are as follows: 1. Hybrid/Manual Decomposition Technique. A method for effectively identifying trends and seasonal components within cloud performance traces. 2. Automatic Decomposition Technique. A fully automated approach for decomposing cloud performance traces without human intervention, which provides decomposition results nearly equivalent to the hybrid method. 3. Comprehensive Case Study. An analysis of diverse cloud traces, revealing that seemingly random and irregular traces still exhibit regularity and predictability, enabling more effective performance and resource management in the cloud. 4. Two applications of decomposition – Decompositionbased predictions and resource optimization that not only serve as a practical application of our decomposition techniques but also validate their correctness and benefits. The rest of this paper is organized as follows: Section II and Section III introduce time-series decomposition and the SAB serverless application; Section IV and Section V present the hybrid and automatic decomposition techniques, which are compared in Section VI; Section VII evaluates both decomposition techniques using other SAB functions; Section VIII presents decomposition-informed resource optimization; Section IX presents related work; Section X discusses limitations; and Section XI concludes this paper. II. T IME - SERIES D ECOMPOSITION M ETHODS Time-series decomposition is a statistical method that breaks a time series into distinct components. As shown in equation (1) [28], [29], a series typically consists of a trend (systematic changes in the mean), seasonal cycles (e.g., quarterly or monthly patterns), other cyclic variations, and random noise. Note that, depending on the data, these components may combine multiplicatively rather than additively in Eq (1). time series = trend + seasonalities + cycles + noise (1)
Latency (ms)
3
CaptureStripe
550 a) 500 450 550 b) 500 450 c) 0 25 25 d) 0
B. Fourier Transform Fourier transform (FT) [31] is a widely used time-series decomposition technique based on frequency analysis. FT can decompose a signal into its discrete frequency components, such as distinguishing different sound frequencies in an audio recording. However, FT struggles with non-stationary signals [32], making it unsuitable for analyzing non-stationary cloud performance traces.
Trend Seasonal Residual
C. Ensemble Empirical Mode Decomposition
07/25/ 09/13/ 11/03/ 12/24/ 02/12/ 04/03/ 05/23/ 20 20 20 20 21 21 21 Fig. 2: STL decomposition on the trace from Fig. 1. CaptureStripe Upper Envelope Lower Envelope IMF1 Average of Lower and Upper Envelopes
Latency (ms)
475 450 10 0 10 08/14/
20
08/24/ 20
09/03/ 20
09/13/ 20
09/23/
20
Fig. 3: An illustration of one iteration of EMD decomposition using a portion of the trace from Fig. 1.
A key characteristic of a time series is “stationarity.” A time series is “stationary” if it exhibits no systematic changes in the mean (i.e., no trend) and variance, and has strictly periodic variations removed [28]. Analyzing non-stationary time series often requires more advanced statistical techniques. Due to multi-tenancy, long-term cloud performance data are typically non-stationary, as exemplified by Fig. 1.
A. STL Decomposition STL [20] decomposition employs “LOESS smoothing” to extract both trend and seasonal components. Due to its simplicity, STL has been widely used in performance anomaly detection for cloud and distributed computing [17], [18], [19]. However, it has been reported that anomaly detection using STL may underperform compared to other methods [21]. Our findings align with this report, as we also observed that applying STL naively to cloud performance traces struggles to accurately identify all components. Fig. 2 presents the STL decomposition of the performance trace from Fig. 1. As shown, STL extracts only a single seasonal component with a bi-weekly cycle, while the monthly and quarterly cycles remain undetected. Additionally, a Runs Test [30] confirms that the STL residual is not random, indicating the presence of unresolved components.
Unlike traditional decomposition techniques, data-driven decomposition imposes minimal prior assumptions, such as stationarity, making them well-suited for handling “irregular” time series [32]. One of the most widely used data-driven methods is Empirical Mode Decomposition (EMD) [26], which is purely empirical and places few constraints on the data [33], [34], [35]. The components extracted through EMD are known as Intrinsic Mode Functions (IMFs). EMD is based on the Hilbert-Huang Transform. Fig. 3 illustrates how this transform is applied to the trace from Fig. 1 to generate one IMF (component). The process begins by identifying local maxima and minima in the trace. Next, cubic spline interpolation is used to construct upper and lower envelopes based on these extrema (i.e., the red and blue lines in Fig. 3). The averages of these envelopes are then computed (shown as the green line in Fig. 3). Finally, the first IMF (IMF1 in Fig. 3) is obtained by subtracting the envelope averages from the original trace. Intuitively, the envelope averages capture the primary fluctuation (variation) in the trace. The difference between the original trace and this primary fluctuation represents the smaller deviations, which form a decomposed component, i.e., an IMF. After removing an IMF from the trace, the remaining residual can undergo further decomposition using the HilbertHuang Transform. By iteratively applying this process, additional variation components (IMFs) are extracted until the final residual reaches a low standard deviation [26]. A key limitation of EMD is the aforementioned “mode mixing” issue, where a single IMF may contain multiple small and/or intermittent components [27], [36]. To address this issue, Ensemble EMD (EEMD) was introduced [27]. EEMD enhances EMD by adding white noise to the original trace, perturbing the signal to facilitate the identification of small intermittent variations. Beyond EMD/EEMD, other data-driven decomposition techniques exist [32]. As EEMD performed the best among them, we adopted it in this study. D. Hybrid/manual Decomposition Manually decomposing a trace by selecting different, more suitable mathematical models for each component is also common [37], [38], [39], [40]. For instance, trend identification can be performed using methods like moving average, linear regression, piecewise linear regression, or LOESS regression [37]. Seasonal components can be extracted using various cyclic models, such as sinusoidal regression, ARIMA [37], or
4
Start
Search Flights
Create Stripe
Process Booking
Reserve Booking
Collect Payment
Get Loyalty
List Bookings
Ingest Loyalty
Notify Booking
Confirm Booking
Capture Stripe
Fig. 4: The invocation chain of the SAB functions. exponential smoothing [41]. This hybrid modeling approach is particularly effective for decomposing cloud performance traces, which exhibit diverse behaviors. Furthermore, hybrid decomposition also allows adjusting of model parameters. For instance, if a trace has intermittent weekly cycles, determining the cycle period can be difficult for automatic fitting due to this intermittency. However, a hybrid decomposition can use a cyclic model with a fixed 7-day period to model the weekly cycles. By manually tuning parameters, the decomposition becomes more robust against signal intermittency, reducing the risk of skewed results. The process of selecting models and their parameters also enhances the reasoning of performance fluctuations. For example, if a decomposed component successfully captures a weekly cycle, it clearly indicates that the performance fluctuation is driven by the differences among weekdays. This need for explainability is also why we avoid using neural networks (NN) in our decomposition. III. C ASE S TUDY DATA S ET As a case study, we applied our decomposition techniques to the year-long performance traces from the Serverless Airline Booking Application (SAB) [16]. Developed by AWS, SAB serves as a representative cloud-native application (CNA), showcasing best practices for building CNAs using AWS services. It provides comprehensive coverage of AWS cloudnative services, such as cloud-native databases (DynamoDB), event-driven workflows (AWS Step Functions), content delivery (CloudFront), messaging and notifications (AWS SNS), and instrumentation (CloudWatch). Fig. 4 illustrates all 11 functions within SAB along with their invocation chain. These functions encompass a standard flight booking process, including flight searches, booking transactions, and user account management. Note that, a SAB function may be invoked by itself (i.e., out of the chain) depending on the use case. Fig. 5 presents the performance traces of SAB functions, collected from July 2020 to June 2021. Specifically, for each function, daily performance data under 100 simultaneous invocations were recorded over 329 days (47 weeks). Our decompositions were applied to the first 301 days (43 weeks), while the remaining 28 days (4 weeks) served as test datasets to demonstrate the predictive capabilities of our decomposition, as well as its evaluation. As shown in Fig. 5, the 11 functions exhibit diverse behaviors and distinct patterns of performance fluctuations. Successfully decomposing all of them highlights the versatility of our techniques. It is worth noting that some traces may appear similar due to shared underlying services. For instance, CollectPayment, CreateStripe, and CaptureStripe show
comparable patterns as they all rely on the payment service from stripe.com. Similarly, ConfirmBooking and ReserveBooking exhibit similar trends due to their use of the same DynamoDB. However, despite these similarities, each trace retains unique characteristics–such as randomness, outliers, trends, and seasonality–necessitating individual decomposition processes (also as illustrated later by the decomposed components with different periods shown in Table I). IV. H YBRID /M ANUAL D ECOMPOSITION A. Overview of Hybrid/Manual Decomposition Fig. 6 outlines the seven-step workflow of our manual decomposition process. Step 1 processes the input performance trace to remove outliers. Here, we use the Hampel Filter [42], although most outlier-remove algorithms also work. Step 2 involves identifying the long-term trend using a regression model. Users can select a regression model suited to the trace’s trajectory. In our case study, we primarily used linear and piece-wise linear regression, though alternatives like moving averages or LOESS can also be applied. Step 3 focuses on detecting quarterly or bi-quarterly cycles, if present. Based on our experience, sinusoidal regression is most effective for these cycles, though any cyclic models can be used, such as ARIMA or Holt-Winters Exponential Smoothing (HWES) [43]. Step 4 identifies (bi-)monthly cycles with cyclic models. Note that, in our case study, we found some traces lack clear monthly cycles. Hence, cyclic model parameters must be carefully tuned in this step. Step 5 captures (bi-)weekly cycles, for which we typically use HWES. Step 6 identifies any remaining cyclic patterns. In our case study, we usually can find semiweekly cycles with HWES regression. Finally, Step 7 applies the Runs Randomness Test [30] to the residual. If the residual is deemed random, the decomposition is complete; otherwise, Step 6 is repeated. B. Hybrid/Manual Decomposition Example To illustrate the process of the hybrid decomposition, we apply it to the performance trace of the SAB function CaptureStripe from Fig. 1. The decomposition results are presented in Fig. 7, and the detailed cycle periods of each component are given in Table I. Following the steps outlined in Fig. 6, we began by preprocessing and identifying the long-term trend (Steps 1&2). Noting the steady increase in latency, we hypothesized that the mean latency followed a linear trend [28]. Consequently, we applied linear regression to model this trend, as shown in Fig. 7a). If this hypothesis holds, the observed linear increase could be attributed to hardware or software changes within the cloud infrastructure [44], [45], [13]. Or, more plausibly, it may result from a progressively increasing background load – CaptureStripe invokes web services from stripe.com, an online payment service. It is likely that stripe.com experienced a steady rise in workload from 2020 [46], causing CaptureStripe to exhibit a corresponding linear increase in latency. For Step 3, which involves identifying quarterly cycles, we analyzed the residual latency trace after removing the linear trend (i.e., the difference between the original latency
5
Latency (ms)
600 500
CollectPayment
600 CreateStripe 400 200 GetLoyalty 100 35.0 IngestLoyalty 32.5 200 ListBookings 100 07/25 09/13 11/03 12/24 02/12 04/03 05/23 /20 /20 /20 /20 /21 /21 /21
40 30 300 200 100 200 100
NotifyBooking
17.5 15.0
ReserveBooking
ProcessBooking SearchFlights
17.5 ConfirmBooking 15.0 07/25 09/13 11/03 12/24 02/12 04/03 05/23 /20 /20 /20 /20 /21 /21 /21
Fig. 5: Latency traces of SAB functions (the trace of CaptureStripe is in Fig. 1). Performance Trace
7.Random -ness test
Yes, is random, stop decomposition
3. Identify (bi-)quarterly cycles, if any
2. Identify long-term trend
1. Preprocess
6. Identify other cycles, if any
4. Identify (multi-)monthly cycles, if any
5. Identify (multi-)weekly cycles, if any.
No, not random, repeat
Fig. 6: The overall workflow of manual decomposition.
and linear regression). As shown in Fig. 7b), the residual latency resembles a sine wave, indicating bi-quarterly variations. The first lowest point appeared around mid-December 2020, aligning with the common expectation that data center usage declines during the holiday season, leading to lower resource contention and improved latency. Observations that align with real-world expectations serve as circumstantial evidence supporting the validity of a decomposition [28], [47]. Fig. 7b) also shows a peak around March/April 2021, followed by a decline in May 2021. Seasonal fluctuations in web traffic (high in spring and low in summer) are common [48], [49], [50]. Higher demand in spring can increase resource contention, resulting in higher latency, while lower demand in summer can reduce contention and improve performance. Given the sine-shaped pattern in Fig. 7b), we fitted a sine function with a 180-day period to represent this bi-quarterly cyclic component. Fig. 7c) presents the residual latency after removing the bi-quarterly variation. To identify potential monthly patterns (Steps 4), we performed weekly averaging on the residuals, which revealed multiple sinusoidal patterns. Consequently, we applied two sine regressions and identified a bi-monthly component (about 56 days) and a semi-monthly (15-day) component. These two components were combined and are shown in Fig. 7c). Fig. 7d) presents the residues after removing all previously identified components. The residual exhibits clear weekly cycles, with peak latencies typically occurring on Tuesdays, Wednesdays or Thursdays and lowest latencies on Sundays or
Saturdays. Fig. 8 displays the autocorrelation function (ACF) of the residual. ACF measures how much a time series is correlated with itself at different time lags. Fig. 8 shows the highest correlations are at lags of 7, 14, 21, and 28 days, confirming the presence of weekly cycles. This aligns with the common understanding that data center load and resource contention peak mid-week and decrease over weekend. To model this weekly pattern, we applied HWES regression, which identified a weekly (7-day) cycle and a semi-weekly (4.6-day) cycle, as shown in Fig.7d). After these two HWES regressions, the final residuals were tested using the Runs test [30] and deemed random, marking the completion of the decomposition process. Since a common application of decomposition is prediction, the accuracy of a decomposition can be evaluated based on its prediction performance [28]. Consequently, we combined the decomposed models to forecast the latency for the next 28 days. The predicted latency values and the corresponding observed (ground truth) latency are shown in Fig. 9. We used the first 301 data points as “training data” to decompose, then predicted the latency for the following 28 days with decomposed models. Note that, none of the 28 days’ ground truth latency data were used in the prediction. Fig. 9 demonstrates that the predictions are highly accurate, with an error (Mean Absolute Percentage Error, MAPE) of only 1.5%. More importantly, our predictions successfully captured most peaks and valleys, along with most of the turning points in the observed latency, indicating that the decomposition correctly identified long-term trend and seasonal variations. This high accuracy further confirms the effectiveness of our decomposition technique. V. AUTOMATIC D ECOMPOSITION Our automated decomposition technique is tailored for scenarios that require performance analysis without manual intervention and for users with limited statistical expertise. As previously discussed, this method is built upon EEMD [27].
Latency (ms)
6
a) 550 500 450 Mar/Apr 50 b) 0 Mid Dec 50 50 c) 0 Wed 50 d) 0 50 07/25/ 09/13/ 11/03/ 12/24Sun /20 02/12/21 04/03/21 05/23/21 20 20 20
Original Latency Long Term Trend (linear fit) Residual Latency (trend removed) Quarterly Variation (sine fit) Residual Latency (trend and quarter removed) Semi-Monthly + Bi-Monthly Variation (two sine fit) Residual Latency (all above components removed) Semi-Weekly + Weekly Variation (two HWES fit)
ACF Value
Fig. 7: Trend and seasonal variations decomposed from the trace in Fig. 1 using the hybrid decomposition technique.
1.0 0.5 0.0 0.5
0
7
14
Lags (in days)
21
28
Latency (ms)
Fig. 8: ACF values of the residuals in Fig. 7d).
550 a) Observed Predicted 500 450 07/25/ 09/13/ 11/03/ 12/24/ 02/12/ 04/03/ 05/23/ 20 20 20 20 21 21 21 550 b) Observed Predicted 500 05/24/ 2
1
05/31/ 0 0 0 21 6/07/21 6/14/21 6/21/21
Fig. 9: Latency prediction based on the hybrid decomposition for CaptureStripe.
A. Automatic Decomposition Example Again, we demonstrate the automatic EEMD decomposition using CaptureStripe’s trace from Fig. 1. The decomposition results are presented in Fig. 10, where EEMD breaks down the trace into six components—five Intrinsic Mode Functions (IMFs) and a residue. Notably, EEMD extracts IMFs in descending order of frequency, with each IMF capturing a variation component with a lower frequency. Due to space limitation, the first two components, IMF1 and IMF2 are combined in Fig. 10a). IMF1 (Fig. 10a)), captures the most oscillatory component of the trace, characterized by semi-weekly cycles with a 3.3-day period on average. Similar semi-week components were found for all traces in our case study with both decomposition methods, suggesting this semi-week fluctuation is a common behavior in AWS cloud. The second component, IMF2 (Fig. 10a)) has an average period of 7.6 days. Most of the peaks in IMF2 occur on Tuesdays, Wednesdays, or Thursdays, while most valleys are observed on Saturdays or Sundays. That is, IMF2 captures weekly cycles. The third component, IMF3 (Fig. 10b)), corresponds to
semi-monthly cycles with an average period of 18.3 days, while the fourth component, IMF4 (Fig. 10b)), captures bimonthly cycles with an average period of 54 days. These semi-monthly and bi-monthly patterns were also identified by the hybrid decomposition. However, the cycles in IMF3 and IMF4 appear less regular than their counterparts from the hybrid approach, particularly around Nov. and Dec. 2020, where the semi-monthly variations are less visible in the IMF3. This absence of cycles is a case of signal intermittency, which is captured by intermittency-aware EEMD (and the hybrid/manual decomposition). IMF5 (Fig. 10c)) shows notable lows during December and peaks around March, represents the bi-quarterly seasonal fluctuation, which was also found by the hybrid decomposition. The last component, residual (Fig. 10d)), reveals a nearly linearly increasing trend, reflecting consistent growth in latency over time, similar to the linear trend found by the hybrid decomposition. Similar to the hybrid decomposition, we also use the automatic decomposition results to predict the latency for the next 28 days. Fig. 11 illustrates the predicted latency alongside the observed latency. The prediction achieves a low error (MAPE) of just 2.0%. The prediction also correctly captures most peaks and valleys with high fidelity. This precise alignment between predicted and observed values also corroborates that our automatic decomposition effectively identifies the underlying trends and seasonal patterns. VI. C OMPARISON OF THE H YBRID /M ANUAL AND AUTOMATIC D ECOMPOSITION Although Section V showed that both hybrid and automatic decomposition methods identify similar components, differences in amplitude and phase shifts may still exist. Therefore, this section presents a direct comparison of these two techniques using the performance trace from Fig. 1. The decomposed components from both methods are illustrated and compared in Fig. 10. Fig. 10d) compares the long-term trend components identified by both decomposition techniques. As shown, both methods produce nearly identical trends, with the only notable differences occurring near the endpoints, where the trend from the automatic technique appears slightly flattened. This flattening may result from a common issue in EEMD known
Latency (ms)
7
50 a) 0 50 0 b) 25 25 c) 0 25 550 d) 500 450 07/25 09/13 11/03 12/24 02/12 04/03 05/23 /20 /20 /20 /20 /21 /21 /21
Semi-Weekly + Weekly Variation (Hybrid) IMF 1 + IMF 2 Semi-Monthly + Bi-Monthly Variation (Hybrid) IMF 3 + IMF 4 Quarterly Variation (Hybrid) IMF 5 Long Term Trend (Hybrid) Residue
Latency (ms)
Fig. 10: Components (IMFs) from the automatic decomposition, with comparisons with the hybrid/manual decomposition.
550 a) Observed Predicted 500 450 07/25/ 09/13/ 11/03/ 12/24/ 02/12/ 04/03/ 05/23/ 20 20 20 20 21 21 21 550 b) Observed Predicted 500 05/24/ 2
1
05/31/ 0 0 0 21 6/07/21 6/14/21 6/21/21
Fig. 11: Prediction based on Automatic decomposition for CaptureStripe.
as “end effects,” where EEMD may incorrectly decompose near the boundaries of a time series [51], [52]. Nonetheless, to determine whether the difference is really caused by ”end effects” or if the trend genuinely flattens toward the ends, additional data would be required. Figure 10c) compares the quarterly components identified by both techniques. The two components align closely, with only a slight difference at the start of the curve—once again due to the “end effect.” This minor discrepancy causes the automatic decomposition’s quarterly component to have a longer period (221 days) than the hybrid version (180 days). Fig. 10b) compares the semi-monthly and bi-monthly components identified by both techniques. Overall, these components exhibit highly similar patterns, except for the period between November and February. This discrepancy is due to signal intermittency, where these monthly cycles diminished during the holiday season. Being non-parametric, EEMD captured this temporary loss of cycles in its IMF, whereas the hybrid approach assumed more consistent bi-weekly and bimonthly cycles throughout the year. To better understand a cloud application’s performance, accurately capturing cycle loss and intermittency is preferable. However, for performance prediction, intermittency in cyclic patterns can negatively impact modeling accuracy, as seen in the automatic approach’s higher prediction error than the hybrid decomposition (2.0% vs. 1.5% MAPE). Ideally, decomposition should balance both objectives—preserving intermittency while maintaining high prediction accuracy. Achieving this balance, however, requires further research.
Finally, Figure 10a) compares the weekly and semi-weekly components identified by both techniques. While their periods are largely similar, discrepancies appear between November and February due to signal intermittency. Additionally, there are differences in amplitude between the hybrid and automatic decomposition results at the beginning of the traces. These amplitude differences stem from “end effects” and the accumulated differences from other seasonal components. Overall, our hybrid and automatic decomposition techniques yield similar components and show comparable effectiveness across the cloud-native applications studied. Minor differences arise from signal intermittency and ‘end effects,’ which introduce irregular periods and amplitude variations in the automatic results. For anomaly detection, capturing such irregularities can be valuable, whereas for performance prediction, the regularity of the hybrid method is often preferable. In future work, we aim to refine decomposition methods to combine both advantages. VII. R ESULTS OF THE C OMPLETE C ASE S TUDY This section presents decomposition results for the remaining 10 SAB functions. Table I lists component periods, while Fig. 12 shows four examples with the largest (though still minor) differences between hybrid and automatic decompositions. Only four are selected due to space constraints.
A. Decomposed Components 1) Long-term Trend: As shown in Table I, these traces exhibit diverse trends, including upward and downward patterns with potential plateaus, and convex curves that decline before rising. This variety of trends indicates that SAB applications represent a wide range of performance behaviors. Table I also shows that trends from the hybrid and automatic approaches largely agree. The hybrid method mainly uses linear or piecewise linear regression for better explainability, while the non-parametric automatic method produces curved but roughly linear trends (sometimes segmented with breakpoints). Four functions—GetLoyalty, ListBookings, IngestLoyalty, and NotifyBooking exhibit minor differences, likely due to EEMD’s “end-effect,” as shown in Fig. 12.
10 a) 0 10 5 b) 50 5 c) 0
Hybrid
Automatic
Latency (ms)
25 a) 0 5 b) 50 2.5 0.0 c) 2.5 90 d) 85
a) 2.5 0.0 2.5 1 b) 0 0.5 c) 0.0 0.5 32.5 d) 32.0
Latency (ms)
Latency (ms)
Latency (ms)
8
1 a) 10 0.5 0.0 b) 0.5 0 c) 1 26.5 d) Hybrid Automatic 26.0 07/25/ 09/13/ 11/03/ 12/24/ 02/12/ 04/03/ 05/23/ 20 20 20 20 21 21 21
(i) GetLoyalty
35 d) 30 07/25/ 09/13/ 11/03/ 12/24/ 02/12/ 04/03/ 05/23/ 20 20 20 20 21 21 21 (iii) ListBookings
(ii) IngestLoyalty
(iv) NotifyBooking
Fig. 12: Hybrid and automatic decomposition results for four SAB applications. Each application has four figures representing the summed a) weekly components, b) monthly components, c) quarterly components, and d) long-term trend.
2) Seasonal Components: Table I also gives the periods (in days) of each decomposed component from both our hybrid and automatic decomposition methods. Despite the drastically different shapes of the performance traces, their decompositions generally follow weekly, monthly, and quarterly cycles. Specifically, all SAB functions exhibit semi-weekly, weekly, and semi-monthly cycles, while most also display monthly or bi-monthly patterns, along with bi-quarterly cycles. Additionally, a few functions show cycles lasting between 80 to 90 days, roughly corresponding to one season. As mentioned earlier, these periodic variations are likely driven by cyclic changes in the background load (i.e., multitenancy) of the cloud data centers. These background fluctuations lead to variations in hardware resource contention, which in turn cause performance fluctuations. Depending on an application’s resource usage (e.g., compute-intensive vs. memory-intensive), the impact of this resource contention can vary. As a result, the specific seasonal components of SAB functions differ slightly. Additionally, nearly all seasonal components identified by both the methods have similar periods, indicating both detect comparable patterns. However, for CollectPayment and NotifyBooking, the hybrid method revealed a 130-day cycle that EEMD missed, as it was embedded within bi-quarterly components of EEMD. This absence of components reduced prediction accuracy, as shown in TableII, where the hybrid method produced significantly lower errors. For GetLoyalty and ListBookings, the quarterly component periods differ between hybrid and automatic decompositions: 180 days for the hybrid method, versus 223 and 271 days for the automatic method. As with CaptureStripe in Section VI, this discrepancy stems from EEMD’s “end effect,” where IMFs flattening near trace boundaries distorts period estimates.
As shown in Fig.s 12i-c) and 12iii-c), the automatic/EEMD curves flatten at the ends, lengthening the apparent period. Nonetheless, the central portions align well, confirming both methods still identified the same quarterly component. Overall, the seasonal patterns observed indicate that cloud application performance follows clear temporal cycles, suggesting management should account for factors such as weekday and time of year. Moreover, while performance traces may appear random, time-series decomposition reveals valuable patterns. These findings call for further research into applying modern statistical techniques to cloud traces to support more informed, large-scale system management. B. Performance Prediction Results We applied both our decompositions to predict SAB’s performance over next 28 days. As shown in Table II, both methods achieved low average MAPE errors: 1.8% for hybrid and 2.1% for automatic. The hybrid approach performed better, benefiting from human expertise in identifying additional seasonal components, such as the 130-day cycles in CollectPayment and NotifyBooking (Section VII-A2). The largest prediction errors were observed for CreateStripe due to the outliers in the groundtruth, which can be seen with large spike at the beginning of Fig. 13b). When outliers are removed, the prediction errors are only 1.1% and 1.7% for the hybrid and automatic decomposition methods. In addition to the low prediction error, the predicted curves from our decomposition methods accurately capture the peaks and valleys of performance fluctuations, as shown in Fig. 13. The figure also demonstrates that the hybrid decomposition more closely follows the ground truth curve compared to the automatic approach. Overall, these highly accruate predictions
9
Weekly1 4.6 3.3 4.2 3.3 4 3.4 3.7 3.1 4.5 3.2 3.9 3.1 4.3 3 4 3.1 4.5 3.1 4.1 2.8 3 3.1
CaptureStripe (Hybrid) CaptureStripe (Auto) CollectPayment (Hybrid) CollectPayment (Auto) CreateStripe (Hybrid) CreateStripe (Auto) GetLoyalty (Hybrid) GetLoyalty (Auto) ListBookings (Hybrid) ListBookings (Auto) SearchFlight (Hybrid) SearchFlight (Auto) ConfirmBooking (Hybrid) ConfirmBooking (Auto) IngestLoyalty (Hybrid) IngestLoyalty (Auto) ProcessBooking (Hybrid) ProcessBooking (Auto) ReserveBooking (Hybrid) ReserveBooking (Auto) NotifyBooking (Hybrid) NotifyBooking (Auto)
Weekly2 7 7.6 6.4 7.1 6.5 7 5.5 7.2 6.1 6.5 6.1 6.9 5.7 6.8 6 7.2 6.1 7.6 6 7.1 6 6.4
Weekly3 15 18.3 15 17 14 15.3 21 17.1 18 16.8 15 15.2 16 15.9 15 14.3 14.3 16.1 7 15.1 14 14.7
Monthly1 56 54 56 54.2 54.7 44.4 30 35.6 30 37.4 30 30 27.5 30.2 35 35.4 34.9 39.8 28 29.5 59.5 61
Monthly2 N/A N/A N/A N/A N/A N/A 60 76.5 73.7 74.5 72.3 78 70 77 81 80 N/A N/A 68.7 73.5 90 95.5
Monthly3 N/A N/A 136 N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A 129 N/A
Quarterly 180 221 209 205 180 196 180 223 180 271 158 156 N/A N/A N/A N/A 126 117 N/A N/A 209 202
Trend Upward Linear Upward Linear Upward Linear Upward Linear Upward Linear Upward Linear Decline→Plateau, 3-Piece Lin. Downward Linear Convex, 3-Piece Linear Downward Linear Convex, 2-Piece Linear Convex Curve Rise→Plateau, 2-Piece Lin. Rise→Plateau curve Plateau→Decline, 2-Piece Lin. Downward Linear Decline→Plateau, 2-Piece Lin. Decline→Plateau curve Rise→Plateau, 2-Piece Lin. Rise→Plateau curve Plateau→Decline, 2-Piece Lin. Downward Linear
Latency (ms)
TABLE I: Periods (in days) of the decomposed components from the hybrid and automatic decomposition.
600 a) 550 90 c) 85
Observed(Ground Truth) CollectPayment GetLoyalty
Hybrid Predicted Automatic Predicted CreateStripe 700 b) 600 500 33 d) IngestLoyalty 32 31 26 f) NotifyBooking 25
35.0 e) ListBookings 32.5 110 g) ProcessBooking 40.0 h) SearchFlights 37.5 100 15.5 i) ReserveBooking 15.5 j) ConfirmBooking 15.0 15.0 14.5 05/24/21 05/31/21 06/07/21 06/14/21 06/21/21 05/24/21 05/31/21 06/07/21 06/14/21 06/21/21 Fig. 13: Predicted daily latency of the rest 10 SAB traces.
suggest that our decompositions are effective and likely capture all the variation components in cloud performance traces. For comparison, we also applied STL decomposition and LSTM network to predict SAB latency. Table II shows both performed worse than our methods, with STL having the highest average error (3.6%). Moreover, the curves predicted by STL and LSTM frequently mis-predicted the peaks and valleys. 1 They might even predict flat lines, missing all variations. This failure can be quantified using ERP (Edit Distance with Real Penalty) [53], which measures the shape differences of the predicted and actual time series by computing their normalized distance (lower is better). Table II shows that 1 For the clarity of the figures, STL and LSTM curves are omitted from Fig. 13.
LSTM and STL yield much higher ERP values than our methods, reflecting frequent mispredictions of variations. Moreover, unlike our decomposition methods, neither STL nor LSTM can fully report seasonal patterns: LSTM offers limited interpretability as a neural network, while STL extracts only a single cyclic component, as shown in Fig. 2. VIII. A U SE C ASE S TUDY: R ESOURCE A LLOCATION BASED ON D ECOMPOSITION R ESULTS A. Resource Allocation Optimization for SAB Performance decomposition aids not only prediction but also tasks like performance debugging [54] and resource allocation. This section presents a use case applying decomposition insights to optimize cloud resource allocation. Our decomposition results for SAB in Section VII show all SAB functions have clear weekly cycles, with higher (slower)
10
SAB Hybrid Automatic LSTM STL Function (MAPE/ERP) (MAPE/ERP) (MAPE/ERP) (MAPE/ERP) CollectP. 1.3% / 1.3 CreateS. 5.4% / 3.0 CaptureS. 1.5% / 1.4 SearchF. 1.8% / 0.1 ListB. 1.9% / 0.1 1.7% / 0.2 GetL. ConfirmB. 1.1% / 1.0 ReserveB. 0.8% / 0.6 IngestL. 1.3% / 2.1 NotifyB. 1.0% / 0.3 ProcessB. 1.5% / 0.2 Average 1.8% / 0.9
1.7% / 1.7 6.0% / 3.2 2.0% / 1.9 1.9% / 0.1 2.0% / 0.1 1.7% / 0.2 1.4% / 1.3 1.3% / 0.9 1.9% / 2.6 1.6% / 0.4 1.1% / 0.2 2.1% / 1.1
4.5% / 4.5 5.9% / 3.2 2.9% / 2.6 2.3% / 0.1 2.4% / 0.1 2.0% / 0.3 2.0% / 1.9 1.7% / 1.3 2.1% / 3.5 1.3% / 0.4 2.9% / 0.4 2.7% / 1.7
4.5% / 4.5 6.9% / 3.7 3.8% / 3.5 2.4% / 0.1 2.4% / 0.1 4.3% / 0.6 2.5% / 2.3 2.6% / 2.0 2.5% / 3.9 2.0% / 0.6 5.7% / 0.8 3.6% / 2.0
TABLE II: Prediction errors (MAPE and ERP) for SAB. Function
Decomposition Informed
All-Weekday (naive optimization)
Std. Slowest Latency Std. Slowest Latency Reduc. Reduction Reduc. Reduction CollectPayment 55.1% CreateStripe 46.7% CaptureStripe 54.0% SearchFlights 52.6% ListBookings 66.7% GetLoyalty 58.5% ConfirmBooking 78.3% ReserveBooking 66.7% IngestLoyalty 73.4% NotifyBooking 53.6% ProcessBooking 56.8% Average 60.2%
6.0% 10.3% 6.1% 12.7% 7.7% 10.3% 11.9% 12.1% 17.2% 7.6% 10.2% 10.2%
57.0% 40.0% 55.4% 52.6% 57.8% 64.1% 87.0% 71.4% 85.9% 53.6% 57.7% 62.0%
6.0% 10.3% 6.1% 12.7% 7.6% 11.5% 14.1% 12.1% 20.3% 7.6% 10.2% 10.8%
TABLE III: Standard deviation and slowest latency reduction of SAB function latencies under the optimized executions.
latencies on certain weekdays. To improve latency stability and ensure consistent user experience, SAB administrators can choose to optimize resource allocations on the slower weekdays. Therefore, in this use case study, we optimized SAB’s execution on AWS by increasing VM memory allocation to 1024MB memory on each function’s slower weekdays (as indicated by the decomposition results), while maintaining the normal 512MB allocation on the other days. Each SAB function ran with the adjusted resource allocations for two weeks in August 2025 to evaluate its performance (optimized or decomposition-informed execution). For comparison, another set ran concurrently with normal allocations (unoptimized execution). Fig. 14a compares latency under optimized and unoptimized executions. Recall that only the slower weekdays identified by decomposition (marked with gold asterisk) were optimized. As Fig. 14a shows, a SAB function typically had two or three slower weekdays, varying from Monday to Friday. Fig. 14a also shows that latency improved substantially on these days for optimized executions. Overall, optimizing the slower weekdays reduced maximum latency by 10.2% and standard deviation by 60.2% on average (Table III), greatly improving performance and stability. As an additional comparison, we also evaluated naively optimizing all five weekdays. As Table III shows, optimizing
Cloud Benchmark
Decomposition Informed
All-Weekday (naive optimization)
Std. Slowest Lat. Std. Slowest Lat. Reduc. Reduction Reduc. Reduction AWS
IO DB Average
57.9% 60.0% 59.0%
6.3% 15.2% 10.8%
52.3% 70.0% 61.2%
6.9% 16.6% 11.8%
TABLE IV: Standard deviation and slowest latency reduction of Amazon cloud benchmarks under the optimized executions. Cloud
Benchmark
AWS
IO DB
Hybrid (MAPE/ERP)
Automatic (MAPE/ERP)
3.04%/1.5 3.17%/1.6
2.65%/1.4 4.24%/2.2
TABLE V: Prediction errors (MAPE and ERP) of Amazon cloud’s IO and DB benchmarks
all weekdays reduced slowest latency by 10.8% and standard deviation by 62.0%, similar to the gains from decompositioninformed execution. However, as decomposition-informed execution increased resource usages on fewer days, it achieves similar performance at a lower cloud usage cost. Overall, these results show that the insights from decomposition can indeed translate to real benefits for cloud deployments. B. Resource Optimization for Other Benchmarks To show that our decomposition approach extends beyond SAB, we repeated the weekday-optimization experiments on two SeBS benchmarks [55], targeting cloud storage (AWS S3) and database (AWS DynamoDB). The benchmarks were slightly adapted for continuous deployment. These two benchmarks were executed on AWS Cloud for four weeks, with first two weeks’ data decomposed to identify the slower weekdays. The decomposition revealed clear weekly cycles and potential bi-weekly cycles. Prediction errors are shown in Table V. Due to space constraints, we omit detailed decomposition results. We then ran the benchmarks under both optimized (decomposition-informed) and unoptimized configurations for two more weeks. Results are shown in Fig. 14b and Table IV. Decomposition-informed execution reduced slowest latency by 10.8% and standard deviation by 59.0%. Table IV also shows that the performance gain was also comparable to the naive allweekday optimization, but at lower cost due to only selected weekdays were optimized. These findings demonstrate that (1) seasonal variations extend beyond SAB, (2) decomposition works across benchmarks, and (3) decomposition insights translate into tangible performance gains across benchmarks. IX. R ELATED W ORK Twitter applied STL decomposition to cloud performance traces to identify long-term trends before detecting performance anomaly [17]. Similarly, Meta also applied STL decomposition before performance anomaly detection [54]. However, as discussed in Section II-A, STL decomposition cannot
Unoptimized
Decomposition-Informed High-resource day 850 (ii) CollectPayment 800 750 (iv) GetLoyalty 275 250 (v) IngestLoyalty (vi) ListBookings 200 175 (viii) ProcessBooking 400 (ix) SearchFlights 350 (x) ReserveBooking 47.5 45.0 42.5
Sa t Su n Mo n Tue We d Th u Fri Sa t Su n Mo n Tue We d Th u Fri
850 (i) CaptureStripe 800 750 (iii) CreateStripe 1300 1200 1100 100 80 55 50 (vii) NotifyBooking 350 300 45 (xi) ConfirmBooking 40
Sa t Su n Mo n Tue We d Th u Fri Sa t Su n Mo n Tue We d Th u Fri
Latency (ms)
11
(i) AWS IO
Unoptimized
Decomposition-Informed
High-resource day
Fri
u Th
We d
Tue
n Mo
n Su
t Sa
Fri
u Th
We d
Tue
n Mo
Su
n
(ii) AWS DB
t
400 375 14 12
Sa
Latency (ms)
(a) All SAB functions
(b) AWS IO and DB benchmarks
Fig. 14: Two-weeks latency under normal unoptimized vs. decomposition-informed allocation. Gold-asterisk points indicate the weekdays that were executed with higher resources allocation (i.e., optimized days), while the rest days were not optimized.
completely and correctly decompose performance traces from public clouds. Unlike the more controlled environments of internal data centers at Twitter and Meta, public clouds exhibit higher contention and greater randomness, making accurate decomposition more challenging. Besides decomposition, time-series analysis has been applied in fault detection in the HPC and cloud [56], [57], [58], [59], [60]. One work employed ARMA model and Fault Tree Analysis to predict failures [61]. Tavakoli et al. treated log data as windowed time series to schedule straggler jobs [62]. Taerat et al. showed that ARIMA models could predict HPC faults in certain cases [63]. Gottumukkala et al. used the Weibull distribution to model HPC system reliability from time series data [64]. Costa et al. used machine learning to cluster similar HPC data to detect application I/O behaviors and potential variations [65]. Due to the complex nature of fault detection, recent studies tend to employ deep learning rather than just time series analysis [21], [59]. Our work is complementary to these studies as prior work has shown that accurate performance decomposition can improve the accuracy of fault detection by removing seasonality [17], [54]. Similarly, our decomposition techniques can also improve the accuracy of the root causes identification for cloud and HPC performance issues by isolating the impact of individual factors [66], [67], [68], [69], [70], [71], [72], [73], [74], [75], [76], [77], [78], [79], [80].
Another group of closely-related studies involves performance prediction and prediction-based resource management [81], [82], [83], [84], [85], [10], [86]. Liang et al. leveraged multiple machine learning models to predict performance for distributed systems [87]. Fu et al. studied ML-based performance prediction of HPC and cloud, and concluded that the ”inherent variability in performance that fundamentally limits prediction accuracy” [88]. Cherrypick finds the best cloud resource allocation using performance prediction from Bayesian Optimization [89]. Our work differs from these studies as our primary goal is performance decomposition, with prediction serving as a byproduct or an application. That is, rather than forecasting future performance, our work focuses more on identifying the factors or variations that cause cloud performance changes. X. D ISCUSSION AND F UTURE W ORK Our results show time-series decomposition reveals trends and seasonal patterns in seemingly random cloud performance traces. However, broader practical adoption requires further research to test its applicability across more cloud applications and longer traces. This section discusses limitations and future directions. Cloud Applications beyond the Case Study. The SAB application in our case study integrates cloud-native services such as databases and workflows and was tested at a reasonable scale (100 concurrent invocations). However, as SAB
12
mainly represents web services, other workloads—like scientific computing or machine learning—may exhibit different performance characteristics. Thus, further research is needed to evaluate time-series decomposition across a wider range of cloud applications. Multi-year Performance Traces. While the seasonal cycles identified in our case study are supported by predictions and alignment with intuitive patterns (e.g., lower data center contention during holidays), longer performance traces could offer further validation—especially if the observed seasonal cycles consistently appear across multiple years. Moreover, longer performance traces can also alleviate the “end effect” of EEMD decomposition, where EEMD may fail to decompose correctly near the boundaries of a time series. With multi-year traces, these boundary issues—particularly at year-end—can be reduced, enabling more accurate analysis of performance trends across year boundaries and offering deeper insight into the differences between the hybrid and automatic decomposition methods. Other Decomposition Techniques. Time-series decomposition is a rapidly evolving research area. While our hybrid and EEMD-based methods performed well in this case study, cloud applications with different performance characteristics or longer traces may benefit from alternative approaches, such as Variational Mode Decomposition [90] or synchrosqueezed transform [91]. As such, further research should evaluate other decomposition methods, especially when exploring broader applicability of decomposition across other cloud workloads and longer performance traces. XI. C ONCLUSION Time-series decomposition is valuable for cloud performance engineering, but prior STL-based methods struggle with complex public cloud traces. We developed two approaches, hybrid/manual and automatic, and applied them to 11 serverless functions. Both effectively decompose diverse traces, revealing trends and seasonal patterns. Their components enable accurate performance prediction, with average MAPE of 1.8% (hybrid) and 2.1% (automatic). Crucially, our work demonstrates that even a single cloud trace holds rich insights for informed cloud management. ACKNOWLEDGMENT We used OpenAI’s ChatGPT to assist with language editing and refinement of this manuscript. All substantive content, analysis, and conclusions are the authors’ own. R EFERENCES [1] CNCF, “CNCF Cloud Native Definition v1.0,” https://github.com/cncf/toc/blob/main/DEFINITION.md, 2018, accessed: 06-09-2023. [2] Amazon, “What Is Cloud Native?” https://aws.amazon.com/whatis/cloud-native/, 2023, accessed: 06-09-2023. [3] B. Varghese and R. Buyya, “Next Generation Cloud Computing: New Trends and Research Directions,” Future Generation Computer Systems, vol. 79, pp. 849–861, 2018. [4] M. Villamizar, O. Garcés, L. Ochoa, H. Castro, L. Salamanca, M. Verano, R. Casallas, S. Gil, C. Valencia, A. Zambrano, and M. Lang, “Infrastructure Cost Comparison of Running Web Applications in the Cloud Using AWS Lambda and Monolithic and Microservice Architectures,” in 2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid), 2016.
[5] Z. Li, L. Zheng, Y. Zhong, V. Liu, Y. Sheng, X. Jin, Y. Huang, Z. Chen, H. Zhang, J. E. Gonzalez, and I. Stoica, “AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23), 2023. [6] L. Wang, L. Yang, Y. Yu, W. Wang, B. Li, X. Sun, J. He, and L. Zhang, “Morphling: Fast, Near-Optimal Auto-Configuration for Cloud-Native Model Serving,” in Proceedings of the ACM Symposium on Cloud Computing, 2021. [7] Y. Lu, S. Bian, L. Chen, Y. He, Y. Hui, M. Lentz, B. Li, F. Liu, J. Li, Q. Liu, R. Liu, X. Liu, L. Ma, K. Rong, J. Wang, Y. Wu, Y. Wu, H. Zhang, M. Zhang, Q. Zhang, T. Zhou, and D. Zhuo, “Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native,” 2024. [Online]. Available: https://arxiv.org/abs/2401.12230 [8] A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y. Vigfusson, and J. Mace, “Serving DNNs like Clockwork: Performance Predictability from the Bottom Up,” in 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), 2020. [9] A. Janes and B. Russo, “Automatic performance monitoring and regression testing during the transition from monolith to microservices,” in 2019 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW), 2019. [10] S. M. M. Fattah, A. Bouguettaya, and S. Mistry, “Long-term iaas selection using performance discovery,” IEEE Transactions on Services Computing, vol. 15, no. 4, 2022. [11] H. Jayathilaka, C. Krintz, and R. Wolski, “Performance monitoring and root cause analysis for cloud-hosted web applications,” in Proceedings of the 26th International Conference on World Wide Web, 2017. [12] A. Maricq, D. Duplyakin, I. Jimenez, C. Maltzahn, R. Stutsman, and R. Ricci, “Taming Performance Variability,” in USENIX Symp. on Operating Systems Design and Implementation, 2018. [13] K. Figiela, A. Gajek, A. Zima, B. Obrok, and M. Malawski, “Performance Evaluation of Heterogeneous Cloud Functions,” Concurrency and Computation: Practice and Experience, vol. 30, no. 23, p. e4792, 2018. [14] J. Thalheim, A. Rodrigues, I. E. Akkus, P. Bhatotia, R. Chen, B. Viswanath, L. Jiao, and C. Fetzer, “Sieve: Actionable Insights from Monitored Metrics in Distributed Systems,” in Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference, 2017. [15] B. Bugbee, C. Phillips, H. Egan, R. Elmore, K. Gruchalla, and A. Purkayastha, “Prediction and characterization of application power use in a high-performance computing environment,” Statistical Analysis and Data Mining: The ASA Data Science Journal, vol. 10, no. 3, pp. 155–165, 2017. [16] S. Eismann, D. E. Costa, L. Liao, C.-P. Bezemer, W. Shang, A. van Hoorn, and S. Kounev, “A case study on the stability of performance tests for serverless applications,” Journal of Systems and Software, vol. 189, p. 111294, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0164121222000498 [17] O. Vallis, J. Hochenbaum, and A. Kejariwal, “A Novel Technique for Long-Term Anomaly Detection in the Cloud,” in 6th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 14), 2014. [18] S. S, N. S, and S. V. D. K, “Time Series Forecasting of Cloud Resource Usage,” in 2021 IEEE 6th International Conference on Computing, Communication and Automation (ICCCA), 2021. [19] K. Jia, X. Yu, C. Zhang, W. Hu, D. Zhao, and J. Xiang, “Software Aging Prediction for Cloud Services Using a Gate Recurrent Unit Neural Network Model Based on Time Series Decomposition,” IEEE Transactions on Emerging Topics in Computing, 2023. [20] R. B. Cleveland, W. S. Cleveland, J. E. McRae, and I. Terpenning, “STL: A Seasonal-Trend Decomposition,” Journal of Official Statistics, vol. 6, no. 1, pp. 3–73, 1990. [21] S. Schmidl, P. Wenig, and T. Papenbrock, “Anomaly Detection in Time Series: A Comprehensive Evaluation,” Proc. VLDB Endow., vol. 15, no. 9, p. 1779–1797, may 2022. [22] P. B. Weerakody, K. W. Wong, G. Wang, and W. Ela, “A review of irregular time series data handling with gated recurrent neural networks,” Neurocomputing, vol. 441, pp. 161–178, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0925231221003003 [23] C. Klötergens, V. K. Yalavarthi, M. Stubbemann, and L. SchmidtThieme, “Functional latent dynamics for irregularly sampled time series forecasting,” in Machine Learning and Knowledge Discovery in Databases. Research Track, A. Bifet, J. Davis, T. Krilavičius, M. Kull, E. Ntoutsi, and I. Žliobaitė, Eds., 2024.
13
[24] Y. Gao, G. Ge, Z. Sheng, and E. Sang, “Analysis and solution to the mode mixing phenomenon in emd,” in 2008 Congress on Image and Signal Processing, vol. 5, 2008, pp. 223–227. [25] G. Xu, Z. Yang, and S. Wang, “Study on mode mixing problem of empirical mode decomposition,” in 2016 Joint International Information Technology, Mechanical and Electronic Engineering Conference. Atlantis Press, 2016, pp. 389–394. [26] N. E. Huang, Z. Shen, S. R. Long, M. C. Wu, H. H. Shih, Q. Zheng, N.C. Yen, C. C. Tung, and H. H. Liu, “The Empirical Mode Decomposition and the Hilbert Spectrum for Nonlinear and Non-stationary Time Series Analysis,” Proceedings of the Royal Society of London. Series A: mathematical, physical and engineering sciences, vol. 454, no. 1971, pp. 903–995, 1998. [27] Z. Wu and N. E. Huang, “Ensemble empirical mode decomposition: a noise-assisted data analysis method,” Advances in Adaptive Data Analysis, vol. 01, no. 01, pp. 1–41, 2009. [28] C. Chatfield and H. Xing, The analysis of time series: an introduction with R. CRC press, 2019. [29] P. J. Brockwell and R. A. Davis, Introduction to Time Series and Forecasting. Springer, 2002. [30] P. Massoli, “Exploring time series randomness,” Curr Res Stat Math, vol. 3, no. 1, pp. 01–07, 2024. [31] R. N. Bracewell, “The Fourier Transform,” Scientific American, vol. 260, no. 6, pp. 86–95, 1989. [Online]. Available: http://www.jstor.org/stable/24987290 [32] T. Eriksen and N. u. Rehman, “Data-driven Nonstationary Signal Decomposition Approaches: A Comparative Analysis,” Scientific Reports, vol. 13, no. 1, p. 1798, 2023. [33] S. Motamedi-Fakhr, M. Moshrefi-Torbati, M. Hill, C. M. Hill, and P. R. White, “Signal Processing Techniques Applied to Human Sleep EEG Signals—A Review,” Biomedical Signal Processing and Control, vol. 10, pp. 21–33, 2014. [34] Y.-H. Wang, K. Hu, and M.-T. Lo, “Uniform Phase Empirical Mode Decomposition: An Optimal Hybridization of Masking Signal and Ensemble Approaches,” IEEE Access, vol. 6, 2018. [35] A. Stallone, A. Cicone, and M. Materassi, “New Insights and Best Practices for the Successful Use of Empirical Mode Decomposition, Iterative Filtering and Derived Algorithms,” Scientific reports, vol. 10, no. 1, p. 15161, 2020. [36] N. E. Huang, Z. Shen, and S. R. Long, “A New View of Nonlinear Water Waves: the Hilbert Spectrum,” Annual review of fluid mechanics, vol. 31, no. 1, pp. 417–457, 1999. [37] R. J. Hyndman and G. Athanasopoulos, Forecasting: principles and practice. OTexts, 2018. [38] S.-X. Lv and L. Wang, “Deep learning combined wind speed forecasting with hybrid time series decomposition and multi-objective parameter optimization,” Applied Energy, vol. 311, p. 118674, 2022. [39] J. F. de Oliveira and T. B. Ludermir, “A Hybrid Evolutionary Decomposition System for Time Series Forecasting,” Neurocomputing, vol. 180, pp. 27–34, 2016, progress in Intelligent Systems Design. [40] G. Zhang, “Time Series Forecasting using A Hybrid ARIMA and Neural Network Model,” Neurocomputing, vol. 50, pp. 159–175, 2003. [41] E. S. Gardner Jr, “Exponential Smoothing: The State of the Art,” Journal of forecasting, vol. 4, no. 1, pp. 1–28, 1985. [42] R. K. Pearson, Y. Neuvo, J. Astola, and M. Gabbouj, “Generalized hampel filters,” EURASIP Journal on Advances in Signal Processing, vol. 2016, pp. 1–18, 2016. [43] P. R. Winters, “Forecasting Sales by Exponentially Weighted Moving Averages,” Management science, vol. 6, no. 3, pp. 324–342, 1960. [44] A. Uta, A. Custura, D. Duplyakin, I. Jimenez, J. Rellermeyer, C. Maltzahn, R. Ricci, and A. Iosup, “Is Big Data Performance Reproducible in Modern Cloud Networks? ,” in 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20), 2020. [45] H. S. Gunawi, R. O. Suminto, R. Sears, C. Golliher, S. Sundararaman, X. Lin, T. Emami, W. Sheng, N. Bidokhti, C. McCaffrey, D. Srinivasan, B. Panda, A. Baptist, G. Grider, P. M. Fields, K. Harms, R. B. Ross, A. Jacobson, R. Ricci, K. Webb, P. Alvaro, H. B. Runesha, M. Hao, and H. Li, “Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems,” ACM Trans. Storage, vol. 14, no. 3, oct 2018. [46] Stripe, “Stripe 2021 Update,” https://stripe.com/files/stripe-2021update.pdf, 2021, [Online]. [47] E. Ates, B. Aksar, V. J. Leung, and A. K. Coskun, “Counterfactual explanations for multivariate time series,” in 2021 International Conference on Applied Artificial Intelligence (ICAPAI), 2021.
[48] K. Pratt, “How Your Web Traffic Changes With the Season,” https://uberall.com/en-us/resources/blog/how-your-web-traffic-changeswith-the-season, 2019, accessed: 07-09-2023. [49] M. Technologies, “Notice a drop in your Website Traffic this summer? Don’t Panic!” https://www.mltinnovations.com/notice-a-drop-inyour-website-traffic-this-summer-dont-panic/, 2023, accessed: 07-092023. [50] S. Pace, “How Does Summer Affect Website Traffic?” https://blog.imageworksllc.com/blog/how-does-summer-affect-websitetraffic, 2015, accessed: 07-09-2023. [51] T. Xiong, Y. Bao, and Z. Hu, “Does Restraining End Effect Matter in EMD-based Modeling Framework for Time Series Prediction? Some Experimental Evidences,” Neurocomputing, vol. 123, pp. 174–184, 2014, contains Special issue articles: Advances in Pattern Recognition Applications and Methods. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0925231213007340 [52] Y. Lei, Z. He, and Y. Zi, “Application of the EEMD Method to Rotor Fault Diagnosis of Rotating Machinery,” Mechanical Systems and Signal Processing, vol. 23, no. 4, pp. 1327–1338, 2009. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0888327008002720 [53] L. Chen and R. Ng, “On the Marriage of Lp-norms and Edit Distance,” in Proceedings of the Thirtieth International Conference on Very Large Data Bases - Volume 30, 2004. [54] D. Y. Yoon, Y. Wang, M. Yu, E. Huang, J. I. Jones, A. Kukkadapu, O. Kocas, J. Wiepert, K. Goenka, S. Chen, Y. Lin, Z. Huang, J. Kong, M. Chow, and C. Tang, “FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production Monitoring,” in Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, 2024. [55] M. Copik, G. Kwasniewski, M. Besta, M. Podstawski, and T. Hoefler, “SeBS: A Serverless Benchmark Suite for Function-as-a-Service Computing,” in Proceedings of the 22nd International Middleware Conference, ser. Middleware ’21. Association for Computing Machinery, 2021. [Online]. Available: https://doi.org/10.1145/3464298.3476133 [56] D. Jauk, D. Yang, and M. Schulz, “Predicting Faults in High Performance Computing Systems: An in-Depth Survey of the State-ofthe-Practice,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2019. [57] A. L. d. C. D. Lima, V. M. Aranha, C. J. a. d. L. Carvalho, and E. G. S. Nascimento, “Smart Predictive Maintenance for HighPerformance Computing Systems: A Literature Review,” J. Supercomput., vol. 77, no. 11, p. 13494–13513, nov 2021. [58] T. Hagemann and K. Katsarou, “A Systematic Review on Anomaly Detection for Cloud Computing Environments,” in Proceedings of the 2020 3rd Artificial Intelligence and Cloud Computing Conference, 2021. [59] H. Ren, B. Xu, Y. Wang, C. Yi, C. Huang, X. Kou, T. Xing, M. Yang, J. Tong, and Q. Zhang, “Time-Series Anomaly Detection Service at Microsoft,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019. [60] O. Ibidunmoye, F. Hernández-Rodriguez, and E. Elmroth, “Performance anomaly detection and bottleneck identification,” ACM Comput. Surv., vol. 48, no. 1, jul 2015. [61] T. Chalermarrewong, T. Achalakul, and S. C. W. See, “Failure prediction of data centers using time series and fault tree analysis,” in 2012 IEEE 18th International Conference on Parallel and Distributed Systems, 2012. [62] N. Tavakoli, D. Dai, and Y. Chen, “Log-Assisted Straggler-Aware I/O Scheduler for High-End Computing,” in 2016 45th International Conference on Parallel Processing Workshops (ICPPW), 2016. [63] N. Taerat, C. Leangsuksun, C. Chandler, and N. Naksinehaboon, “Proficiency Metrics for Failure Prediction in High Performance Computing,” in International Symposium on Parallel and Distributed Processing with Applications, 2010. [64] N. R. Gottumukkala, R. Nassar, M. Paun, C. B. Leangsuksun, and S. L. Scott, “Reliability of a System of k Nodes for High Performance Computing Applications,” IEEE Transactions on Reliability, vol. 59, no. 1, pp. 162–169, 2010. [65] E. Costa, T. Patel, B. Schwaller, J. M. Brandt, and D. Tiwari, “Systematically Inferring I/O Performance Variability by Examining Repetitive Job Behavior,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2021. [66] Y. Gan, Y. Zhang, K. Hu, D. Cheng, Y. He, M. Pancholi, and C. Delimitrou, “Seer: Leveraging Big Data to Navigate the Complexity of Performance Debugging in Cloud Microservices,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, 2019.
14
[67] M. Ma, J. Xu, Y. Wang, P. Chen, Z. Zhang, and P. Wang, “AutoMAP: Diagnose Your Microservice-Based Web Applications Automatically,” in Proceedings of The Web Conference 2020, 2020. [68] M. Li, Z. Li, K. Yin, X. Nie, W. Zhang, K. Sui, and D. Pei, “Causal inference-based root cause analysis for online service systems with intervention recognition,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022. [69] Z. Li, N. Zhao, M. Li, X. Lu, L. Wang, D. Chang, X. Nie, L. Cao, W. Zhang, K. Sui, Y. Wang, X. Du, G. Duan, and D. Pei, “Actionable and Interpretable Fault Localization for Recurring Failures in Online Service Systems,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022. [70] Y. Gan, M. Liang, S. Dev, D. Lo, and C. Delimitrou, “Enabling Practical Cloud Performance Debugging with Unsupervised Learning,” SIGOPS Oper. Syst. Rev., vol. 56, no. 1, p. 34–41, jun 2022. [71] T. Inagaki, Y. Ueda, M. Ohara, S. Choochotkaew, M. Amaral, S. Trent, T. Chiba, and Q. Zhang, “Detecting Layered Bottlenecks in Microservices,” in 2022 IEEE 15th International Conference on Cloud Computing (CLOUD), 2022, pp. 385–396. [72] J. Rios, S. Jha, and L. Shwartz, “Localizing and Explaining Faults in Microservices Using Distributed Tracing,” in 2022 IEEE 15th International Conference on Cloud Computing (CLOUD), 2022, pp. 489–499. [73] R. Xie, J. Yang, J. Li, and L. Wang, “ImpactTracer: Root Cause Localization in Microservices Based on Fault Propagation Modeling,” in 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2023. [74] Z. Zhang, M. K. Ramanathan, P. Raj, A. Parwal, T. Sherwood, and M. Chabbi, “CRISP: Critical Path Analysis of Large-Scale Microservice Architectures,” in 2022 USENIX Annual Technical Conference (USENIX ATC 22), 2022. [75] M. K. Aguilera, J. C. Mogul, J. L. Wiener, P. Reynolds, and A. Muthitacharoen, “Performance Debugging for Distributed Systems of Black Boxes,” SIGOPS Oper. Syst. Rev., vol. 37, no. 5, p. 74–89, oct 2003. [76] I. Cohen, J. S. Chase, M. Goldszmidt, T. Kelly, and J. Symons, “Correlating Instrumentation Data to System States: A Building Block for Automated Diagnosis and Control,” in 6th Symposium on Operating Systems Design & Implementation (OSDI 04), 2004. [77] K. Nagaraj, C. Killian, and J. Neville, “Structured comparative analysis of systems logs to diagnose performance problems,” in Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation, 2012. [78] R. K, P. Tammana, P. G. Kannan, and P. Naik, “A Case For CrossDomain Observability to Debug Performance Issues in Microservices,” in 2022 IEEE 15th International Conference on Cloud Computing (CLOUD), 2022. [79] X. Kong, Y. Zhu, H. Zhou, Z. Jiang, J. Ye, C. Guo, and D. Zhuo, “Col-
lie: Finding Performance Anomalies in RDMA Subsystems,” in 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22), 2022. [80] K. Huck, X. Wu, A. Dubey, A. Georgiadou, J. A. Harris, T. Klosterman, M. Trappett, and K. Weide, “Performance Debugging and Tuning of Flash-X with Data Analysis Tools,” in 2022 IEEE/ACM Workshop on Programming and Performance Visualization Tools (ProTools), 2022. [81] C.-J. M. Liang, H. Xue, M. Yang, L. Zhou, L. Zhu, Z. L. Li, Z. Wang, Q. Chen, Q. Zhang, C. Liu, and W. Dai, “AutoSys: The Design and Operation of Learning-Augmented Systems,” in 2020 USENIX Annual Technical Conference (USENIX ATC 20), 2020. [82] D. Van Aken, A. Pavlo, G. J. Gordon, and B. Zhang, “Automatic Database Management System Tuning Through Large-Scale Machine Learning,” in Proceedings of the 2017 ACM International Conference on Management of Data, 2017. [83] Z. L. Li, C.-J. M. Liang, W. He, L. Zhu, W. Dai, J. Jiang, and G. Sun, “Metis: Robustly Tuning Tail Latencies of Cloud Systems,” in 2018 USENIX Annual Technical Conference (USENIX ATC 18), 2018. [84] J. Cheng, C. Gao, and Z. Zheng, “HINNPerf: Hierarchical Interaction Neural Network for Performance Prediction of Configurable Systems,” ACM Trans. Softw. Eng. Methodol., vol. 32, no. 2, mar 2023. [85] S. Eismann, L. Bui, J. Grohmann, C. Abad, N. Herbst, and S. Kounev, “Sizeless: Predicting the Optimal Size of Serverless Functions,” in Proceedings of the 22nd International Middleware Conference, 2021. [86] C. Delimitrou and C. Kozyrakis, “Paragon: QoS-Aware Scheduling for Heterogeneous Datacenters,” in Proceedings of the Eighteenth International Conference on Architectural Support for Programming Languages and Operating Systems, 2013. [87] C.-J. M. Liang, Z. Fang, Y. Xie, F. Yang, Z. L. Li, L. L. Zhang, M. Yang, and L. Zhou, “On Modular Learning of Distributed Systems for Predicting End-to-End Latency,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), 2023. [88] S. Fu, S. Gupta, R. Mittal, and S. Ratnasamy, “On the use of ML for blackbox system performance prediction,” in 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21), 2021. [89] O. Alipourfard, H. H. Liu, J. Chen, S. Venkataraman, M. Yu, and M. Zhang, “CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics,” in USENIX Symp. on Networked Systems Design and Implementation, 2017. [90] K. Dragomiretskiy and D. Zosso, “Variational Mode Decomposition,” IEEE Transactions on Signal Processing, vol. 62, no. 3, pp. 531–544, 2014. [91] G. Thakur, E. Brevdo, N. S. Fučkar, and H.-T. Wu, “The Synchrosqueezing algorithm for time-varying spectral analysis: Robustness properties and new paleoclimate applications,” Signal Processing, vol. 93, no. 5, pp. 1079–1094, 2013. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0165168412004240