arXiv:2606.19255v1 [cs.LG] 17 Jun 2026
SCAN: Enhance Time Series Anomaly Detection via Multi-Scale Neighborhood-Centered Clustering
Xingze Zheng1 , Hanyin Cheng1 , Siyuan Wang1 , Yiting Hao1 , Peng Chen2 , Yuan Jun3 , Yang Shu1∗ 1 East China Normal University 2 2012 APPLab, Huawei 3 Huawei {xzzheng,hycheng,sywang,ythao}@stu.ecnu.edu.cn, {chenpeng192,yuanjun25}@huawei.com, {yshu}@dase.ecnu.edu.cn
Abstract Time series anomaly detection plays a crucial role in a wide range of real-world applications. Reconstruction-based methods have become the mainstream paradigm, but they suffer from over-generalization and under-generalization problems, which are challenging to balance. To address this, we introduce multi-scale clustering to enhance reconstruction-based methods. At the representation level, we integrate the cluster center representations of normal patterns to constrain the model to target representative normal patterns for reconstruction, preventing dominance of powerful capacity and representation capability. At the anomaly criterion level, we derive anomaly confidence score based on cluster membership probability and combine it with reconstruction error, providing dual criteria for detection. Furthermore, the effectiveness of the cluster center representations and anomaly confidence score depends on the clustering performance. Accordingly, we extract neighborhood-centered representations for multi-view clustering to improve clustering performance. Extensive experiments on multiple real-world datasets from diverse application domains demonstrate the state-of-the-art performance of SCAN.
1
Introduction
With the rapid advancement of digitalization, time series analysis has been widely applied across diverse sectors, including finance, transportation, healthcare, meteorology, energy, and the Internet of Things [1–3]. Mining historical time series data reveals intrinsic temporal patterns and latent regularities, providing reliable decision support for downstream tasks and practical decision-making [4–6]. As a key branch, time series anomaly detection identifies outlier points or anomalous subsequences that deviate from normal patterns [7], enabling early detection of potential anomalies and timely implementation of preventive measures to mitigate risks and avoid heavy losses. Time series anomaly detection is categorized into supervised and unsupervised anomaly detection based on labeled data requirements [8, 9]. In practice, anomalous data is scarce and diverse, rendering its collection and labeling highly challenging. Thus, unsupervised anomaly detection has become a widely studied and applied practical solution [10], which identifies anomalies solely by learning from normal data. Among these methods, reconstruction-based methods become the mainstream paradigm [11]. However, reconstruction-based methods suffer from over-generalization and undergeneralization issues, as shown in Figure 1. Over-generalized model reconstructs simple anomalies as well as normal ones, causing high false negative rates, while under-generalized one fails to reconstruct complex normal patterns, leading to high false positive rates. In both cases, the reconstruction errors of normal and anomalous data lose their distinctiveness, making it impossible to detect anomalies. To ∗ Corresponding author
Preprint.
mitigate over-generalization and under-generalization issues in reconstruction-based methods, two key research challenges remain:
Figure 1: Example of over-generalization and under-generalization. From top to bottom: anomalous series, over-generalized reconstructed series, and under-generalized one. Over-generalization reconstructs both normal and anomalous patterns, while under-generalization fails to reconstruct complex normal patterns. First, traditional reconstruction paradigms fail to balance over-generalization and undergeneralization. Learning normal sample representations and detecting anomalies via reconstruction error inherently involves a trade-off: powerful capacity and representation capability causes overgeneralization [12], while inadequate representation learning leads to under-generalization. A single reconstruction dimension cannot concurrently balance suppressing simple anomaly missdetection by narrowing the generalization boundary and reducing complex normal sample false detection by enhancing representation capability, thus limiting detection accuracy in complex temporal scenarios. To address this challenge, we innovatively propose a novel anomaly detection paradigm by integrating clustering into reconstruction-based methods at the representation and anomaly criterion levels. This is due to clustering can alleviate the above contradiction from the sample distribution perspective [13]. On one hand, it can capture representative normal patterns to constrain the scope of the model’s reconstruction objectives and thus suppress over-generalization. On the other hand, its cluster membership probabilities serve as another criterion for anomaly detection, thereby avoiding complete reliance on reconstruction-based results. Specifically, at the representation level, we first capture rich normal patterns via multi-scale modeling since single-scale modeling captures only limited normal patterns [11]. Then, each scale generates cluster-weighted representations for its time series through weighted aggregation of all cluster center representations, followed by inter-scale fusion of these cluster-weighted representations and intra-scale fusion of the latter with original representations. This constrains the model to target representative normal patterns for reconstruction, thus suppressing over-generalization. At the anomaly criterion level, we derive anomaly confidence score from cluster membership probability and combine it with reconstruction error to form a dual detection criterion. Second, the effectiveness of clustering enhancement strategy depends on cluster partition quality, although integrating clustering with reconstruction is a promising solution to Challenge 1. Shared background context across different temporal sub-patterns (e.g., basic fluctuation trends of industrial sensor data under different operating conditions) [14] impair accurate sample distribution partitioning, resulting in poor intra-cluster compactness and ambiguous inter-cluster boundaries, as shown in Figure 2a. Low-quality clustering fails to provide reliable support for anomaly detection. To address this challenge, we design a clustering module enhanced by neighborhood-centered representations. Neighborhood-centered representations indicate the degree of deviation between temporal patch representations and their neighborhood-weighted aggregated representations, which project patch-level information into a more discriminative space, as shown in Figure 2b. Based on this, the differences among temporal sub-patterns can be amplified by multi-view clustering. We first extract two types of neighborhood-centered representations (based on similarity and temporal dependence respectively) from original representations as auxiliary clustering views, and perform normal pattern clustering on each view to obtain cluster centers. Then, we dynamically filter high2
quality clustering pseudo-labels from neighborhood-centered representation views and use them as supervisory signals to enhance the clustering of original representations. Normal Pattern 1
Normal Pattern 2
(Based on Similarity) (a) Original Rrepresentation Space
Normal Pattern 3
(Based on Temporal Dependence)
(b) Neighborhood-Centered Representation Space
Figure 2: Comparison of original and neighborhood-centered representation spaces. (a) Normal pattern partitioning in the original representation space, where inter-pattern boundaries are ambiguous; (b) Normal pattern partitioning in the neighborhood-centered representation space, where inter-pattern coverage is significantly reduced. All spaces are learned by SCAN on the SMAP dataset. In summary, our contributions in this paper are as follows: • We propose a novel anomaly detection paradigm that integrates multi-scale clustering into traditional reconstruction-based methods at both the representation and anomaly criterion levels, effectively alleviating the inherent over-generalization and under-generalization issues of reconstruction-based methods. • We design a neighborhood-centered clustering module that extracts neighborhood-centered representations to provide trusted supervision for the clustering of original representations, thus furnishing reliable support for anomaly detection. • Our method outperforms state-of-the-art models and achieves superior performance on diverse datasets.
2
Related Work
Time Series Anomaly Detection. Time series anomaly detection methods are classified into traditional and deep learning methods [15]. Traditional methods, including LOF [16], COF [17], OCSVM [18], and SVDD [19], detect anomalies through boundary definition or density estimation, yet fail to effectively model complex temporal dependencies. Deep learning methods adopt contrastive [20, 21], forecasting [22–25], or reconstruction paradigms. Recently, reconstruction-based methods [26–31], such as TimesNet [32], ModernTCN [33], and CrossAD [11], have shown great promise. Despite these advances, conventional reconstruction-based methods still suffer from a critical trade-off between over-generalization and under-generalization. Memory-based Methods. Memory architectures augment neural networks with external storage to retain and retrieve long-term information, with applications spanning natural language processing [34–36], computer vision [37, 38], and other domains [39–42]. Recently, memory-based methods have also been introduced in anomaly detection [43–46], especially in the field of computer vision, such as MemAE [47] and MNAD [48]. For time series, MEMTO [49] demonstrates the effectiveness of leveraging external memory to capture diverse normal patterns across heterogeneous domains. Nevertheless, most methods lack a reliable mechanism to amplify the differences between different temporal patterns.
3
Methodology
Problem Formulation. Given a multivariate time series X ∈ RT ×C with length T and C channels, the task of time series anomaly detection is to learn representations from normal training series and output labels Ŷtest = (ŷ1 , ŷ2 , . . . , ŷT ′ ) for unseen test series Xtest of length T ′ . Each label ŷt ∈ {0, 1} indicates whether the t-th observation xt is normal or anomalous. Generally, 0 indicates normal, while 1 indicates anomalous. 3
a Overall Network
����
����
������
ℒ���
Head
Hierarchical Cluster Fusion
NeighborhoodCentered Clustering
Patch Embedding
RevIN & Patching
Down Sample
ℒ��� ℒ��� ℒ��� ℒ��� ���� ���� ������
Reconstruction Loss Clustering Loss Cross-entropy Constraint Loss Consistency Constraint Loss Reconstruction Error Anomaly Confidence Anomaly Score
c Hierarchical Cluster Fusion G3
b Neighborhood-Centered Clustering (displayed with a single scale) Q
⊗
Temporal Matrix
G5 Selected Label Set
K&V
⊖
Original Representation
K&V
Label Filter
Projection
⊖
Q
Cross Attention
Similarity Matrix
Projection
⊗
Q
G4
ℒ���
ℒ���
K&V
Fused Representation
Training and Inference Pathways
Training-only Pathways
Pre-update Cluster Centers
Post-update Cluster Centers
ℒ���
ℒ���
Neighborhood-Centered Representation
G1
G2
⊕�
⊕�
⊕�
Scale 1
Scale 2
Scale 3
Reconstructed Time Series
The elements within them are related to the original representation. Cluster Membership Probabilities
⊕�
Weighted Aggregation
G
Fusion Gate
Figure 3: The architecture of SCAN, taking scale number j = 3 as an example. 3.1
Overall Architecture
The overall architecture of SCAN is shown in Figure 3. We first adopt multi-scale modeling to capture diverse normal patterns, including scale-specific ones: taking the raw time series X as Xj , we apply j non-overlapping average pooling operations with varying temporal kernel sizes to generate series with j different scales {X0 , X1 , . . . , Xj }, where the length Ti of Xi ∈ RTi ×C satisfies Ti < Ti+1 . Each Xi undergoes reversible instance normalization, is split into Ni patches [50], and is projected into a d-dimensional space to obtain the embedding Ei ∈ RNi ×d×C . These embeddings are fed into the Neighborhood-Centered Clustering Module to extract neighborhood-centered representations. Cross-attention is then applied to cluster both neighborhood-centered and original representations, yielding normal pattern cluster centers and corresponding cluster membership probabilities. The neighborhood-centered branches generate high-quality clustering pseudo-labels through label filtering, which supervise the clustering of the original representation branch. Subsequently, in the Hierarchical Cluster Fusion Module, cluster center representations at each scale are aggregated via a membership probability-based gating mechanism to produce scale-specific cluster-weighted representations. These are fused progressively across scales, and further combined with original representations through intra-scale fusion to form the final fused representation Zi ∈ RNi ×d×C . Finally, Zi is mapped back to the input space via a linear projection head to reconstruct the input series. For simplicity, we will use a single-scale example to introduce the method in the following sections. 3.2
Neighborhood-Centered Clustering
To learn discriminative normal pattern cluster centers, we extract neighborhood-centered representations from the original ones. Their clustering processes act as auxiliary branches, offering trusted supervision for the clustering of the original representations. Extraction of Neighborhood-Centered Representation. A neighborhood-centered representation is obtained by a weighted centering operation under the neighborhood context, which captures the deviation of each temporal patch from its neighborhood. Its extraction comprises three steps: adjacency matrix construction, neighborhood representation aggregation, and neighborhood-centered representation computation. First, we simultaneously construct the similarity adjacency matrix Asim ∈ RNi ×Ni and the temporali tim Ni ×Ni dependent adjacency matrix Ai ∈ R to align neighborhood modeling with the data dissim tribution. Ai is computed by weighting flattened embeddings Ẽi ∈ RNi ×(d×C) with learnable 4
weights w ∈ R1×(d×C) to emphasize key features, followed by cosine similarity calculation. Atim is i initialized via patch-level temporal distances and adaptively optimized during training. This process is formalized as: Ẽw i = N (Ẽi ⊙ w),
(1)
w ⊤ Asim = Softmax(Ẽw i i · (Ẽi ) ),
1 Atim [n, m] = , i |n − m| + 1 ′
(2) (3)
′
Atim = Softmax(Atim ), i i
(4)
where Ẽw i denotes the weighted and normalized embedding, N is normalization, ⊙ denotes elementwise multiplication, and |n − m| (0 ≤ |n − m| ≤ Ni − 1) is the temporal distance between the n-th and m-th patches. Based on the two adjacency matrices above, we perform weighted aggregation on the neighbors of each patch to obtain two types of aggregated neighborhood representations, fusing neighborhood information. We then apply centering operations to the embedding Ei using these aggregated representations, yielding the neighborhood-centered representations Esim ∈ RNi ×d×C and Etim ∈ i i Ni ×d×C R : Esim = Ei − Reshape(Asim · Ẽi ), i i
(5)
Etim = Ei − Reshape(Atim · Ẽi ). i i
(6)
Appendix A expounds the superiority of neighborhood-centered representation over original representation. Pattern Clustering. We learn normal pattern cluster centers by clustering the normal patterns in the time series. First, the original representation Ei and neighborhood-centered representations Esim i , Ni ×dr Etim are projected into a d -dimensional clustering space via a linear layer, producing H ∈ R , r i i Ni ×dr tim Ni ×dr Hsim ∈ R , and H ∈ R , respectively. The clustering of each representation is i i independent. For each clustering branch, we initialize a K-cluster center set. Taking the original branch as an example, its cluster centers are denoted as Pi ∈ RK×dr . For a patch representation hin drawn from Hi , its cluster membership probability uin,ik for the k-th cluster pik is defined as the normalized cosine similarity between hin and pik , formulated as: p⊤ ik hin uin,ik = N ∈ [0, 1], (7) ∥pik ∥∥hin ∥ PK where k=1 uin,ik = 1. Then, we extract normal pattern information via the masked cross-attention mechanism [51]: ⊤
H) b i = N (exp( (WQ Pi )(W √ K i ) ⊙ M⊤ P i )WV Hi , d
(8)
b i denotes the updated cluster center representation, and Mi ∈ RNi ×K is the binary mask where P matrix, i.e., the binary cluster membership matrix obtained via Gumbel-Softmax Bernoulli sampling, with Min,ik ≈ Bernoulli(uin,ik ). A higher uin,ik leads Min,ik closer to 1. This mask guides crossattention to focus on intra-cluster normal patterns, and thus guarantees the separability of normal pattern clusters. Moreover, each clustering branch is optimized via a clustering loss Lclu , which aims to maximize intra-cluster similarity and minimize inter-cluster similarity: ⊤ ⊤ Lori (9) clu,i = − tr Mi Si Mi + tr I − Mi Mi Si , m X sim tim Lclu = Lori (10) clu,i + Lclu,i + Lclu,i , i=1 Ni ×Ni
where Si ∈ R is the same-cluster probability matrix constructed using K-means in the clustering representation space. Specifically, we derive the soft assignment Yisoft [n, k] that patch 5
PK n belongs to cluster k via distance-based softmax, and Si [n, m] = k=1 Yisoft [n, k]Yisoft [m, k] measures the probability that n-th patch and m-th patch belong to the same cluster. Trusted Supervision. Clustering of the original representation is prone to interference from intercluster shared background context, leading to a performance bottleneck. To address this, we enhance it using more discriminative clustering views based on neighborhood-centered representations. We further use a filtering mechanism to select high-quality pseudo-labels of the neighborhood-centered representation branches as supervisory signals for the clustering of original representation [52]. In pattern clustering, each branch outputs a clustering result, where each patch corresponds to three cluster membership probability vectors. For the n-th patch, the membership vectors from sim sim the neighborhood-centered representation branches are denoted as usim in = {uin,i1 , . . . , uin,iK } and P (·) K tim tim utim in = {uin,i1 , . . . , uin,iK }, satisfying k=1 uin,ik = 1 with (·) ∈ {sim, tim}. To measure the reliability of cluster assignments, we utilize the properties of Shannon entropy for PK the membership vector u ∈ RK . It is well-established that the entropy H(u) = − k=1 uk log uk reaches its maximum, Hmax = log K, when u follows a uniform distribution, representing the highest uncertainty. Based on this, we define the information redundancy as R(u) = 1 − H(u)/Hmax . To selectively utilize highly certain pseudo-labels, we adopt an entropy-based quality score: Quality(u) = 1 −
H(u) , 2Hmax
(11)
where a higher Quality(u) indicates a more peaked distribution and thus a more reliable cluster assignment. For each neighborhood-centered representation branch, patches are first ranked by their maximum (·) membership probability maxk uin,ik . We then dynamically determine the number of selected pseudojP k sel,(·) (·) Ni labeled patches by summing the quality scores, and choose the top Bi = Quality(u ) in n=1 (·)
patches to form the trusted index set Ωi for the branch. The pseudo-label for each selected patch (·) (·) (·) is defined as c̃in = arg maxk uin,ik , n ∈ Ωi . The trusted pseudo-labels are used to supervise the cluster membership probabilities uin = {uin,i1 , . . . , uin,iK } of the original representation branch. Specifically, we minimize the cross-entropy constraint loss Lent between the original cluster assignments and the filtered pseudo-labels: ! m X 1 X ϕ(c̃in ) log uin , (12) Lent = − |Ωi | i=1 n∈Ωi
where ϕ denotes the one-hot encoding function. Moreover, we use the consistency constraint loss Lcon to enforce consistency between the neighborhood-centered representations and the original representations in the semantic space: Lcon =
m X
(·)
DKL (Ui ∥ Ui ),
(·) ∈ {sim, tim},
(13)
i=1
where Ui ∈ RNi ×K denotes the membership probability matrix. 3.3
Hierarchical Cluster Fusion
To ensure the model target representative normal patterns for reconstruction, maintaining resistance to over-generalization while preserving key details of the original representation, we integrate cluster center representations into the representations used for reconstruction. Inter-scale Fusion. After obtaining the scale-specific cluster centers, we derive a cluster-weighted representation for each patch by weighting the cluster center representations with the corresponding cluster membership probabilities. For the i-th scale, given the updated cluster center representab i ∈ RK×dr and the membership probability matrix Ui ∈ RNi ×K , the cluster-weighted tions P b i. representation Gi ∈ RNi ×dr is computed as Gi = Ui · P 6
To propagate normal patterns from coarser to finer scales, we conduct progressive inter-scale fusion via gating mechanisms. For the coarsest scale, we directly set F0 = G0 . For a finer scale i > 0, we first align the previous fused cluster-weighted representation Fi−1 to the resolution of the current e i−1 , then fuse it with Gi using a learnable gate: scale via linear interpolation to obtain F e i−1 ]), αi = σ(Gateinter [Gi ; F (14) i
e i−1 , Fi = αi ⊙ Gi + (1 − αi ) ⊙ F
(15)
where [·; ·] denotes concatenation and σ(·) denotes the sigmoid function. This design enables the model to adaptively balance scale-specific details and cross-scale global context. Intra-scale Fusion. The fused cluster-weighted representation Fi ∈ RNi ×dr is further integrated with the original patch embedding Ei to construct the final representation for reconstruction. As Fi lies in the cluster representation space while Ei resides in the reconstruction space, we first project Fi to Eclu ∈ RNi ×(d×C) via a linear layer. A learnable gate is then used to fuse Ei and Eclu i i : intra clu β i = σ Gatei [Ei ; Ei ] , (16) Zi = β i ⊙ Ei + (1 − β i ) ⊙ Eclu i ,
(17)
where Zi denotes the final fused representation for reconstruction. By preserving the original embedding and injecting a clustering-guided normal pattern constraint, this intra-scale fusion alleviates both over-generalization to anomalies and under-generalization to normal patterns. Finally, we minimize the Mean Squared Error (MSE) between the input series and its reconstructed result: m X Lrec = ||Xi − X̂i ||22 . (18) i=1
3.4
Loss Function
The objective function of SCAN contains four components: reconstruction loss Lrec , clustering loss Lclu , cross-entropy constraint loss Lent , and consistency constraint loss Lcon . Thus, the total loss Ltotal is defined as follows: Ltotal = Lrec + λ1 Lclu + λ2 Lent + λ3 Lcon ,
(19)
where λ1 , λ2 , and λ3 are the balance coefficients. 3.5
Anomaly Criterion
The anomaly score Stotal of SCAN consists of two components: reconstruction error Srec and anomaly confidence Sclu . We first interpolate the scale-wise anomaly scores with varying lengths to align them to a uniform length Tm : m Y Srec = interpolate((Xi − X̂i )2 ), (20) Sclu =
i=1 m Y
interpolate(1 −
max k∈{1,...,K}
i=1
Stotal = (Srec )1−γ · (Sclu )γ ,
Ui (:, k)),
(21) (22)
where γ is the balancing coefficient and X̂i is the reconstructed results. Following prior works, we run the SPOT [53] to automatically compute the threshold δ after obtaining anomaly score. A point is marked as an anomaly if its anomaly score exceeds δ.
4
Experiments
4.1
Experimental Setup
Datasets. We evaluate on various widely adopted benchmark datasets, including SMD [27], MSL [22], SMAP [22], PSM [54], SWAT [55], and NeurIPS-TS [56] (which encompasses GECCO 7
and SWAN). We also conduct experiments on the UCR dataset [57], with corresponding results presented in Appendix C.2. Baselines. We extensively compare our model with 22 baselines, including the latest state-ofthe-art (SOTA) anomaly detection models. These baselines comprise: linear transformation-based methods: OCSVM [18], PCA [58]; density estimation-based methods: LOF [16], HBOS [59]; outlierbased methods: IForest [60], LODA [61]; neural network-based methods: Autoencoders (AE) [62], DAGMM [26], LSTM [22], OmniAnomaly (Omni) [27], CAE Ensemble (CAE) [63], Anomaly Transformer (AT) [20], TimesNet [32], DC Detector (DC) [21], GPT4TS [64], ModernTCN [33], TimeMixer [65], MtsCID [66], MEMTO [49], DADA [10], CrossAD [11], KAN-AD [25]. Metrics. Many methods refine detection results via point adjustment [21, 20]. However, this strategy incorrectly assumes that a single accurate point detection validates an entire anomalous segment. To address this, some works adopt the Affiliation-F1 [67], yet it is highly threshold-sensitive. Recent studies confirm VUS-PR [68, 69] as the most robust, accurate and fair metric. We thus mainly use VUS-ROC and VUS-PR, alongside mainstream metrics for comprehensive comparison, which are provided in Appendix C.1. Further implementation details are presented in Appendix B. 4.2
Main Results
We evaluate SCAN on seven real-world datasets against 22 baselines, as shown in Table 1. SCAN achieves state-of-the-art performance under both VUS-ROC (V-R) and VUS-PR (V-P), exhibiting superior detection accuracy and stability over various pre-selected thresholds. This validates the effectiveness of our neighborhood-centered clustering and hierarchical cluster fusion, which refine reconstruction representations and produce clustering-based anomaly confidence scores, offering a new paradigm for time series anomaly detection. Table 1: VUS-metrics results in the seven real-world datasets. Higher VUS-ROC (V-R) or VUS-PR (V-P) values indicate better performance. Red: the best, Blue: the 2nd best. Dataset
SMD
MSL
SMAP
PSM
SWaT
GECCO
SWAN
Metric
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
OCSVM PCA IForest LODA HBOS LOF AE DAGMM LSTM CAE Omni AT DC GPT4TS ModernTCN MtsCID TimeMixer TimesNet MEMTO DADA CrossAD KAN-AD
0.6451 0.7174 0.7224 0.6745 0.6670 0.6893 0.7560 0.6988 0.7001 0.7174 0.7080 0.5117 0.5145 0.7679 0.7707 0.5162 0.7711 0.8420 0.5165 0.8276 0.8580 0.7657
0.1131 0.1529 0.1304 0.1213 0.1102 0.1076 0.1542 0.1496 0.1395 0.1376 0.1340 0.0796 0.0814 0.1745 0.1596 0.0815 0.1391 0.2040 0.0824 0.1624 0.2344 0.1593
0.5798 0.6108 0.5638 0.5375 0.6265 0.6081 0.6047 0.6069 0.6163 0.5382 0.5490 0.3890 0.3900 0.7697 0.7747 0.4686 0.7858 0.7880 0.5077 0.7918 0.8091 0.6629
0.1753 0.1889 0.1631 0.1689 0.1790 0.1715 0.1890 0.1803 0.1681 0.1639 0.1973 0.1041 0.0948 0.2769 0.3010 0.1181 0.2461 0.2731 0.1521 0.3028 0.3144 0.2107
0.4185 0.4090 0.4960 0.3973 0.5620 0.5673 0.4687 0.5599 0.5329 0.4212 0.4743 0.4571 0.4444 0.5449 0.5470 0.4260 0.5552 0.5495 0.5059 0.4535 0.5779 0.4772
0.1133 0.1144 0.1315 0.1017 0.1388 0.1409 0.1366 0.1349 0.1399 0.1140 0.1239 0.1239 0.1149 0.1289 0.1395 0.1177 0.1371 0.1352 0.1416 0.1223 0.1443 0.1276
0.5993 0.6331 0.6009 0.6089 0.7056 0.6628 0.6339 0.5598 0.5571 0.6113 0.6340 0.5186 0.5235 0.6466 0.6480 0.5194 0.5974 0.6344 0.5208 0.6940 0.7302 0.6377
0.4252 0.4706 0.3964 0.4423 0.5061 0.4615 0.4490 0.4522 0.4592 0.4395 0.4472 0.3309 0.3366 0.4599 0.4668 0.3296 0.3807 0.4373 0.3307 0.4782 0.5596 0.4523
0.5903 0.6149 0.3677 0.6358 0.7084 0.6667 0.5903 0.5746 0.5482 0.5939 0.6187 0.5561 0.5191 0.2537 0.2735 0.5021 0.2673 0.2974 0.7354 0.7714 0.7865 0.7821
0.4396 0.4459 0.1011 0.3531 0.4602 0.4187 0.4144 0.4731 0.2200 0.4104 0.4475 0.2679 0.1495 0.0846 0.0941 0.1283 0.0918 0.1158 0.2858 0.4422 0.4767 0.5829
0.7533 0.5366 0.7083 0.5749 0.5440 0.7817 0.6124 0.5099 0.6450 0.5524 0.5386 0.4751 0.5454 0.9776 0.9694 0.5315 0.9899 0.9834 0.6018 0.9791 0.9948 0.8567
0.1207 0.0443 0.0943 0.0339 0.0453 0.0919 0.0448 0.0396 0.0668 0.0528 0.0517 0.0278 0.0361 0.4181 0.4819 0.0381 0.4606 0.4578 0.0449 0.4588 0.6211 0.1732
0.9088 0.9290 0.8835 0.9170 0.9056 0.9095 0.6982 0.8951 0.9082 0.9042 0.9041 0.8046 0.8429 0.9340 0.9027 0.8128 0.9290 0.9515 0.8158 0.9521 0.9499 0.4433
0.9004 0.9123 0.8793 0.9107 0.8894 0.9007 0.0201 0.8697 0.8862 0.9022 0.9022 0.7943 0.8338 0.8924 0.8962 0.8375 0.8721 0.9160 0.8389 0.9124 0.9171 0.3909
SCAN
0.9081
0.3203
0.8232
0.3315
0.6114
0.1578
0.7457
0.5851
0.7866
0.4986
0.9957
0.6812
0.9591
0.9233
4.3
Model Analysis
Ablation Studies. To validate the effectiveness of SCAN, we conduct ablation studies on its key modules, as shown in Table 2. Multi-scale modeling consistently improves model performance. Based on this, pattern clustering enables reconstruction with cluster center representations, further boosting metrics. Applying trusted supervision to pattern clustering increases VUS-ROC and VUSPR by 1.49% and 9.42%, respectively, owing to effective anomaly confidence scores. Integrating cluster center representations into original features improves the two metrics by 2.01% and 8.63%, respectively, demonstrating the necessity of representation fusion. Moreover, removing the multiscale cluster fusion module degrades performance, verifying such fusion enriches normal pattern modeling. 8
Table 2: Ablations on the key components of SCAN, including Multi-Scale Modeling, Pattern Clustering, Trusted Supervision, and Cluster Fusion. Higher VUS-ROC (V-R) or VUS-PR (V-P) values indicate better performance. Bold: the best. Dataset
PSM
GECCO
Avg.
Multi-Scale Modeling
Pattern Clustering
Trusted Supervision
Cluster Fusion
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
✗ ✓ ✓ ✓ ✗ ✓
✗ ✗ ✓ ✓ ✓ ✓
✗ ✗ ✗ ✓ ✓ ✓
✗ ✗ ✗ ✗ ✗ ✓
0.6754 0.6901 0.7008 0.7185 0.7090 0.7457
0.4638 0.4936 0.5118 0.5416 0.5278 0.5851
0.5067 0.5348 0.5876 0.5963 0.5767 0.6114
0.1312 0.1375 0.1495 0.1528 0.1477 0.1578
0.9767 0.9817 0.9842 0.9916 0.9879 0.9957
0.4542 0.5096 0.5368 0.6167 0.5742 0.6812
0.7196 0.7356 0.7575 0.7688 0.7579 0.7843
0.3497 0.3802 0.3994 0.4370 0.4165 0.4747
Model Efficiency. We compare the efficiency of SCAN with representative time series anomaly detection methods, including MLP-based (TimeMixer), CNN-based (ModernTCN, TimesNet, KAN-AD) and Transformer-based (AnomalyTransformer, CrossAD) methods, using both static and runtime metrics, as shown in Table 4. SCAN achieves superior detection performance alongside speed advantage. In addition, the time of SCAN P complexity m 2 is O i=1 Ni (demb + K)+ Ni (demb dr + Kdr + Ldemb ) with controllable overhead, where demb denotes the embedding dimension (demb = d × C), and L represents the patch size. The detailed complexity analysis is presented in Appendix C.3. Visualization. To validate the efficacy of clustering for anomaly detection, we perform a comparative visualization of clustering-based and traditional anomaly scores, as shown in Figure 5. The time series contains two complex anomalies: shapelet and seasonal. Traditional criterion exhibits high false positive and negative rates, while clustering-based criterion, which integrates reconstruction error and anomaly confidence scores, enables stable and accurate detection.
5
SMAP
Figure 4: Static and runtime performance metrics of SCAN and other time series anomaly detection methods on the GECCO dataset, evaluated with train epochs of 5 and a batch size of 128. Method
Training Time (s) Inference Time (s) Total Params (M)
ModernTCN TimesNet TimeMixer AnomalyTransformer CrossAD KAN-AD
22.75 342.79 57.2 270.49 270.75 68.92
0.57 1.68 0.87 15.05 0.80 0.12
0.05 4.68 0.1 4.74 0.93 0.41
SCAN
15.11
0.15
1.76
Figure 5: Visualization of traditional and clustering-based anomaly scores. From top to bottom: time series, traditional anomaly scores, clustering-based anomaly scores. Red regions mark anomalies.
Conclusion
In this work, we propose SCAN, a clustering-enhanced time series anomaly detection model. By integrating clustering and reconstruction within a novel paradigm, and leveraging neighborhood-centered representations to boost clustering, we mitigate the over-generalization and under-generalization trade-off dilemma in reconstruction-based methods. However, like most clustering methods, SCAN requires a predefined number of clusters, which limits its flexibility. In future work, we plan to evaluate more real-world scenarios and explore adaptive clustering algorithms.
9
References [1] T. Zhou, Z. Ma, Q. Wen, X. Wang, L. Sun, and R. Jin, “Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting,” in Proceedings of the International Conference on Machine Learning (ICML), pp. 27268–27286, 2022. [2] P. Chen, Y. ZHANG, Y. Cheng, Y. Shu, Y. Wang, Q. Wen, B. Yang, and C. Guo, “Pathformer: Multi-scale transformers with adaptive pathways for time series forecasting,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024. [3] Y. Wang, Y. Qiu, P. Chen, Y. Shu, Z. Rao, L. Pan, B. Yang, and C. Guo, “Lightgts: A lightweight general time series forecasting model,” in Proceedings of the International Conference on Machine Learning (ICML), 2025. [4] H. Wu, J. Xu, J. Wang, and M. Long, “Autoformer: Decomposition transformers with Auto-Correlation for long-term series forecasting,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 34, pp. 22419–22430, 2021. [5] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A time series is worth 64 words: Long-term forecasting with transformers,” in Proceedings of the International Conference on Learning Representations (ICLR), 2023. [6] Y. Chen, S. Huang, Y. Cheng, P. Chen, Z. Rao, Y. Shu, B. Yang, L. Pan, and C. Guo, “Aimts: Augmented series and image contrastive learning for time series classification,” in Proceedings of the International Conference on Data Engineering (ICDE), pp. 1952–1965, 2025. [7] X. Wu, X. Qiu, Z. Li, Y. Wang, J. Hu, C. Guo, H. Xiong, and B. Yang, “Catch: Channel-aware multivariate time series anomaly detection via frequency patching,” in Proceedings of the International Conference on Learning Representations (ICLR), 2025. [8] Y. Wang, H. Wu, J. Dong, Y. Liu, C. Wang, M. Long, and J. Wang, “Deep time series models: A comprehensive survey and benchmark,” arXiv preprint arXiv:2407.13278, 2024. [9] X. Qiu, Z. Li, W. Qiu, S. Hu, L. Zhou, X. Wu, Z. Li, C. Guo, A. Zhou, Z. Sheng, J. Hu, C. S. Jensen, and B. Yang, “Tab: Unified benchmarking of time series anomaly detection methods,” in Proceedings of the VLDB Endowment (VLDB), 2025. [10] Q. Shentu, B. Li, K. Zhao, Y. Shu, Z. Rao, L. Pan, B. Yang, and C. Guo, “Towards a general time series anomaly detector with adaptive bottlenecks and dual adversarial decoders,” in Proceedings of the International Conference on Learning Representations (ICLR), 2025. [11] B. Li, Q. Shentu, Y. Shu, H. Zhang, M. Li, N. Jin, B. Yang, and C. Guo, “Crossad: Time series anomaly detection with cross-scale associations and cross-window modeling,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. [12] S. Yoon, J. Kim, J. Ha, and Y. M. Ko, “Momemto: Patch-based memory gate model in time series foundation model,” arXiv preprint arXiv:2509.18751, 2025. [13] X. Qiu, X. Wu, Y. Lin, C. Guo, J. Hu, and B. Yang, “Duet: Dual clustering enhanced multivariate time series forecasting,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 1185–1196, 2025. [14] Z. Tan, X. Luo, Y. Liu, and Y. Zhang, “Mask the redundancy: Evolving masking representation learning for multivariate time-series clustering,” arXiv preprint arXiv:2511.17008, 2025. [15] Y. Zhao, L. Deng, X. Chen, C. Guo, B. Yang, T. Kieu, F. Huang, T. B. Pedersen, K. Zheng, and C. S. Jensen, “A comparative study on unsupervised anomaly detection for time series: Experiments and analysis,” arXiv preprint arXiv:2209.04635, 2022. [16] M. M. Breunig, H.-P. Kriegel, R. T. Ng, and J. Sander, “Lof: identifying density-based local outliers,” in Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD), pp. 93–104, 2000. [17] J. Tang, Z. Chen, A. W.-C. Fu, and D. W. Cheung, “Enhancing effectiveness of outlier detections for low density patterns,” in Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), pp. 535–548, 2002. [18] B. Schölkopf, R. C. Williamson, A. Smola, J. Shawe-Taylor, and J. Platt, “Support vector method for novelty detection,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 12, 1999.
10
[19] D. M. J. Tax and R. P. W. Duin, “Support vector data description,” Machine Learning, vol. 54, no. 1, pp. 45–66, 2004. [20] J. Xu, H. Wu, J. Wang, and M. Long, “Anomaly transformer: Time series anomaly detection with association discrepancy,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022. [21] Y. Yang, C. Zhang, T. Zhou, Q. Wen, and L. Sun, “Dcdetector: Dual attention contrastive representation learning for time series anomaly detection,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 3033–3045, 2023. [22] K. Hundman, V. Constantinou, C. Laporte, I. Colwell, and T. Söderström, “Detecting spacecraft anomalies using lstms and nonparametric dynamic thresholding,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 387–395, 2018. [23] M. Munir, S. A. Siddiqui, A. Dengel, and S. Ahmed, “Deepant: A deep learning approach for unsupervised anomaly detection in time series,” IEEE Access, vol. 7, pp. 1991–2005, 2018. [24] A. Deng and B. Hooi, “Graph neural network-based anomaly detection in multivariate time series,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 35, pp. 4027–4035, 2021. [25] Q. Zhou, C. Pei, F. Sun, Z. Gao, H. Zhang, G. Xie, D. Pei, et al., “Kan-ad: Time series anomaly detection with kolmogorov–arnold networks,” in Proceedings of the International Conference on Machine Learning (ICML), 2025. [26] B. Zong, Q. Song, M. R. Min, W. Cheng, C. Lumezanu, D. Cho, and H. Chen, “Deep autoencoding gaussian mixture model for unsupervised anomaly detection,” in Proceedings of the International Conference on Learning Representations (ICLR), 2018. [27] Y. Su, Y. Zhao, C. Niu, R. Liu, W. Sun, and D. Pei, “Robust anomaly detection for multivariate time series through stochastic recurrent neural network,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 2828–2837, 2019. [28] B. Zhou, S. Liu, B. Hooi, X. Cheng, and J. Ye, “Beatgan: Anomalous rhythm detection using adversarially generated time series,” in Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), vol. 2019, pp. 4433–4439, 2019. [29] J. Audibert, P. Michiardi, F. Guyard, S. Marti, and M. A. Zuluaga, “Usad: Unsupervised anomaly detection on multivariate time series,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 3395–3404, 2020. [30] Z. Li, Y. Zhao, J. Han, Y. Su, R. Jiao, X. Wen, and D. Pei, “Multivariate time series anomaly detection and interpretation using hierarchical inter-metric and temporal embedding,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 3220–3230, 2021. [31] S. Tuli, G. Casale, and N. R. Jennings, “Tranad: Deep transformer networks for anomaly detection in multivariate time series data,” in Proceedings of the VLDB Endowment (VLDB), vol. 15, pp. 1201–1214, 2022. [32] H. Wu, T. Hu, Y. Liu, H. Zhou, J. Wang, and M. Long, “Timesnet: Temporal 2d-variation modeling for general time series analysis,” in Proceedings of the International Conference on Learning Representations (ICLR), 2023. [33] D. Luo and X. Wang, “Moderntcn: A modern pure convolution structure for general time series analysis,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024. [34] G. Lample, A. Sablayrolles, M. Ranzato, L. Denoyer, and H. Jégou, “Large memory layers with product keys,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 32, pp. 8546–8557, 2019. [35] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 33, pp. 9459–9474, 2020. [36] W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei, “Augmenting language models with long-term memory,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 36, pp. 74530–74543, 2023.
11
[37] J. Lei, L. Wang, Y. Shen, D. Yu, T. Berg, and M. Bansal, “Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), pp. 2603–2614, 2020. [38] S. W. Oh, J.-Y. Lee, N. Xu, and S. J. Kim, “Video object segmentation using space-time memory networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9225–9234, 2019. [39] H. Le, T. Karimpanal George, M. Abdolshah, T. Tran, and S. Venkatesh, “Model-based episodic memory induces dynamic hybrid controls,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 34, pp. 30313–30325, 2021. [40] J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 30, pp. 4077–4087, 2017. [41] K. Tanwisuth, X. Fan, H. Zheng, S. Zhang, H. Zhang, B. Chen, and M. Zhou, “A prototype-oriented framework for unsupervised domain adaptation,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 34, pp. 17194–17208, 2021. [42] D. Guo, L. Tian, M. Zhang, M. Zhou, and H. Zha, “Learning prototype-oriented set representations for meta-learning,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022. [43] H. Zhou, J. Yu, and W. Yang, “Dual memory units with uncertainty regulation for weakly supervised video anomaly detection,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 37, pp. 3769–3777, 2023. [44] Y. Lai, Y. Han, and Y. Wang, “Anomaly detection with prototype-guided discriminative latent embeddings,” in Proceedings of the IEEE International Conference on Data Mining (ICDM), pp. 300–309, 2021. [45] R. de Paula Monteiro, M. C. Lozada, D. R. C. Mendieta, R. V. S. Loja, and C. J. A. Bastos Filho, “A hybrid prototype selection-based deep learning approach for anomaly detection in industrial machines,” Expert Systems with Applications, vol. 204, p. 117528, 2022. [46] J. Liu, K. Song, M. Feng, Y. Yan, Z. Tu, and L. Zhu, “Semi-supervised anomaly detection with dual prototypes autoencoder for industrial surface inspection,” Optics and Lasers in Engineering, vol. 136, p. 106324, 2021. [47] D. Gong, L. Liu, V. Le, B. Saha, M. R. Mansour, S. Venkatesh, and A. v. d. Hengel, “Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1705–1714, 2019. [48] H. Park, J. Noh, and B. Ham, “Learning memory-guided normality for anomaly detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14360–14369, 2020. [49] J. Song, K. Kim, J. Oh, and S. Cho, “Memto: Memory-guided transformer for multivariate time series anomaly detection,” Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 36, pp. 57947–57963, 2023. [50] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A time series is worth 64 words: Long-term forecasting with transformers,” in Proceedings of the International Conference on Learning Representations (ICLR), 2023. [51] J. Chen, J. E. Lenssen, A. Feng, W. Hu, M. Fey, L. Tassiulas, J. Leskovec, and R. Ying, “From similarity to superiority: Channel clustering for time series forecasting,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 37, pp. 130635–130663, 2024. [52] Y. Li, M. Yang, D. Peng, T. Li, J. Huang, and X. Peng, “Twin contrastive learning for online clustering,” International Journal of Computer Vision, vol. 130, no. 9, pp. 2205–2221, 2022. [53] A. Siffer, P.-A. Fouque, A. Termier, and C. Largouet, “Anomaly detection in streams with extreme value theory,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 1067–1075, 2017. [54] A. Abdulaal, Z. Liu, and T. Lancewicki, “Practical approach to asynchronous multivariate time series anomaly detection and localization,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 2485–2494, 2021.
12
[55] A. P. Mathur and N. O. Tippenhauer, “Swat: A water treatment testbed for research and training on ics security,” in Proceedings of the International Workshop on Cyber-physical Systems for Smart Water Networks (CySWater), pp. 31–36, 2016. [56] K.-H. Lai, D. Zha, J. Xu, Y. Zhao, G. Wang, and X. Hu, “Revisiting time series outlier detection: Definitions and benchmarks,” in Proceedings of the Annual Conference on Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Track on Datasets and Benchmarks), 2021. [57] R. Wu and E. J. Keogh, “Current time series anomaly detection benchmarks are flawed and are creating the illusion of progress,” IEEE transactions on knowledge and data engineering (TKDE), vol. 35, no. 3, pp. 2421–2429, 2021. [58] M.-L. Shyu, S.-C. Chen, K. Sarinnapakorn, and L. Chang, “A novel anomaly detection scheme based on principal component classifier,” in Proceedings of the IEEE International Conference on Data Mining (ICDM), 2003. [59] M. Goldstein and A. Dengel, “Histogram-based outlier score (hbos): A fast unsupervised anomaly detection algorithm,” in KI-2012: Poster and Demo Track, pp. 59–63, 2012. [60] F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation forest,” in Proceedings of the IEEE International Conference on Data Engineering (ICDE), pp. 413–422, 2008. [61] T. Pevnỳ, “Loda: Lightweight on-line detector of anomalies,” Machine Learning, vol. 102, no. 2, pp. 275– 304, 2016. [62] M. Sakurada and T. Yairi, “Anomaly detection using autoencoders with nonlinear dimensionality reduction,” in Proceedings of the MLSDA Workshop on Machine Learning for Sensory Data Analysis, p. 4–11, 2014. [63] D. Campos, T. Kieu, C. Guo, F. Huang, K. Zheng, B. Yang, and C. S. Jensen, “Unsupervised time series outlier detection with diversity-driven convolutional ensembles,” in Proceedings of the VLDB Endowment (VLDB), vol. 15, pp. 611–623, 2021. [64] T. Zhou, P. Niu, L. Sun, R. Jin, et al., “One fits all: Power general time series analysis by pretrained lm,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), vol. 36, pp. 43322–43355, 2023. [65] S. Wang, H. Wu, X. Shi, T. Hu, H. Luo, L. Ma, J. Y. Zhang, and J. ZHOU, “Timemixer: Decomposable multiscale mixing for time series forecasting,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024. [66] Y. Xie, H. Zhang, and M. A. Babar, “Multivariate time series anomaly detection by capturing coarse-grained intra-and inter-variate dependencies,” in Proceedings of the ACM on Web Conference, pp. 697–705, 2025. [67] A. Huet, J. M. Navarro, and D. Rossi, “Local evaluation of time series anomaly detection algorithms,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD), pp. 635–645, 2022. [68] J. Paparrizos, P. Boniol, T. Palpanas, R. S. Tsay, A. Elmore, and M. J. Franklin, “Volume under the surface: a new accuracy evaluation measure for time-series anomaly detection,” in Proceedings of the VLDB Endowment (VLDB), vol. 15, pp. 2774–2787, 2022. [69] P. Boniol, A. K. Krishna, M. Bruel, Q. Liu, M. Huang, T. Palpanas, R. S. Tsay, A. Elmore, M. J. Franklin, and J. Paparrizos, “Vus: effective and efficient accuracy measures for time-series anomaly detection,” The VLDB Journal, vol. 34, no. 3, p. 32, 2025. [70] X. Qiu, J. Hu, L. Zhou, X. Wu, J. Du, B. Zhang, C. Guo, A. Zhou, C. S. Jensen, Z. Sheng, et al., “Tfb: Towards comprehensive and fair benchmarking of time series forecasting methods,” in Proceedings of the VLDB Endowment (VLDB), vol. 17, pp. 2363–2377, 2024. [71] Y. Liu, H. Zhang, C. Li, X. Huang, J. Wang, and M. Long, “Timer: Generative pre-trained transformers are large time series models,” in Proceedings of the International Conference on Machine Learning (ICML), 2024.
13
A
Analysis: Enhancement of Clustering Separability
To analyze why neighborhood-centered representation can improve clustering, we model each patch embedding as the sum of a specific pattern and a locally shared background context. Based on this model, we state a proposition about the expected change in clustering contrast when we apply neighborhood centering. Definition A.1 (Additive Shared-Context Model). Let the patch embedding ein (the n-th patch at the i-th scale) be expressed as ein = ain + bin , where ain denotes a specific pattern component and bin denotes a locally shared background context component. We define the background-to-pattern ∥bin ∥22 intensity ratio as α = ∥ain (α ≥ 0). ∥2 2
Within a local neighborhood Nin , the background context changes slowly in expectation. Formally, centering by neighborhood averaging mainly reduces the contribution of bin to similarity computations. Definition A.2 (Clustering Contrast). We define the Clustering Contrast ∆ as the expected margin between intra-cluster similarity and inter-cluster similarity: ∆ = E[Sim(ein , eim ) | pin = pim ] − E[Sim(ein , eim ) | pin ̸= pim ],
(23)
where pin and pim denote the clusters of ein and eim relatively, and Sim(·, ·) denotes cosine similarity. A larger ∆ indicates a more separable representation space. Proposition A.3 (Separability Enhancement). Under the Additive Shared-Context Model (Definition A.1), applying neighborhood centering to obtain a neighborhood-centered representation reduces the influence of the shared background context more strongly than it reduces the specific pattern contribution. As a result, the clustering contrast of the neighborhood-centered representation (∆NCR ) is improved compared to that of the original representation (∆OR ), and the improvement is approximately proportional to (1 + α).
B
Implementation Details
B.1
Datasets
In order to verify the effectiveness of our proposed model, we carried out extensive experiments on a range of publicly accessible datasets. These datasets were chosen due to their diverse data characteristics and anomaly types, thereby forming a comprehensive evaluation framework. The selected datasets include SMD (Server Machine Dataset) [27], MSL (Mars Science Laboratory Dataset) [22], SMAP (Soil Moisture Active Passive Dataset) [22], PSM (Pooled Server Metrics Dataset) [54], SWaT (Secure Water Treatment) [55], the GECCO and SWAN sub-datasets from NeurIPS-TS (NeurIPS 2021 Time Series Benchmark) [56], and UCR [57]. Detailed descriptions of each dataset are provided below: 1. SMD collects resource utilization information from the computer clusters owned by an Internet company. 2. MSL, collected by NASA, contains telemetry data that reflects the operational conditions of sensors and actuators on the Martian rover. 3. SMAP, collected by NASA, offers soil moisture data acquired through spacecraft monitoring systems. 4. PSM is sourced from eBay’s server machines and records metrics associated with their operational performance. 5. SWaT includes sensor data from a water treatment infrastructure that operates continuously. 6. NeurIPS-TS is a dataset introduced by [56], and GECCO and SWAN are its sub-datasets, which cover a variety of anomaly scenarios. 7. UCR consists of 250 sub-datasets, with each containing one-dimensional data that has a single anomaly segment. The statistical details of all the aforementioned datasets are summarized in Table 3. 14
Table 3: Statistics of the datasets. The anomaly ratio denotes the abnormal proportion of the entire dataset.
B.2
Dataset
Domain
Dimension
Training
Validation
Test(labeled)
Anomaly Ratio(%)
MSL SMAP PSM SMD SWaT GECCO SWAN UCR
Spacecraft Spacecraft Server Machine Server Machine Water treatment Water treatment Space Weather Natural
1 1 25 38 31 9 38 1
46,653 108,146 105,984 566,724 396,000 55,408 48,000 1,790,680
11,664 27,037 26,497 141,681 99,000 13,852 12,000 447,670
73,729 427,617 87,841 708,420 449,919 69,261 60,000 6,143,541
10.5 12.8 27.8 4.2 12.1 1.25 23.8 0.6
Evaluation Metrics
Time series anomaly detection involves identifying point-based anomalies (individual outliers) and range-based anomalies (continuous outlier segments). Traditional metrics suffer from limitations: threshold-based metrics (e.g., Precision, F1-score) rely on manual threshold tuning; threshold-free metrics (e.g., AUC-ROC, AUC-PR) fail to adapt to range-based anomalies and label misalignments; range-extended metrics (e.g., Range-AUC) still depend on buffer length parameters. Based on recent research [68, 69], the following focuses on the Volume Under the Surface (VUS) series, parameter-free and threshold-free metrics that address these limitations. B.2.1
Preliminaries of Key Metrics • Threshold-based metrics: Discretize anomaly scores via a threshold to compute confusion matrix-based values (e.g., Precision, Recall, F1-score), but are sensitive to threshold choice. • Threshold-free metrics: Compute AUC-ROC/AUC-PR by iterating over all thresholds, but only adapt to point-based anomalies. • Range-AUC series: Extend labels with buffer regions to handle range anomalies, but require tuning the buffer length parameter ℓ.
B.2.2
Parameter-Free Threshold-Free Metrics (VUS Series)
VUS upgrades “area under the curve” to “volume under the 3D surface” by iterating over all buffer lengths and thresholds, achieving fully parameter-free and robust evaluation for time series anomaly detection. Core Foundation: Continuous Label Extension. To handle range anomalies and label misalignments, extend discrete ground-truth labels label ∈ {0, 1}n (0=normal, 1=anomaly) to continuous labels labelℓ ( ℓ = buffer length, default: half the time-series period). For an anomaly interval [s, e]: 1 |s−i| 2 1 − , s − 2ℓ ≤ i < s ∧ predi = 1 ℓ 1, s≤i<e labelℓi = (24) 12 |e−i| ℓ 1 − , e ≤ i < e + ∧ pred = 1 i ℓ 2 0, otherwise where predi is the binary prediction derived from anomaly score ST i and a threshold. If no predicted anomaly exists in the buffer region, labelℓi = 0; overlapping buffers take the maximum value. The components of the extended confusion matrix are defined as follows: the true positive count for class ℓ is denoted as T Pℓ = labelℓ⊤ · pred; the false positive count for class ℓ is expressed as F Pℓ = ⊤
ℓ ) ·I (I − labelℓ )⊤ · pred; the extended positive sample count for class ℓ is given by Pℓ = (label+label , 2 a formulation designed to avoid label distortion; and the range-adapted true positive rate for class ℓ is P i ,P ) computed as T P Rℓ = TPPℓℓ · Ri ∈R ER(R . |R|
VUS-ROC Calculation. Iterate over buffer length set L = [ℓ0 , ℓ1 , ..., ℓL ] ( 0 = ℓ0 < ℓ1 < ... < ℓL = ℓmax ) and threshold set T h = [T h0 , T h1 , ..., T hN ] ( 0 = T h0 < T h1 < ... < T hN = 1 ). 15
The volume under the T P Rℓ -F P Rℓ -ℓ surface is: L N 1 X X (k,w) VUS-ROC = ∆ · |ℓw − ℓw−1 | 4 w=1
(25)
k=1
∆(k,w) = [T P Rℓw (T hk−1 ) + T P Rℓw (T hk )] · [F P Rℓw (T hk ) − F P Rℓw (T hk−1 )] {z } | Contribution of buffer ℓw
+ [T P Rℓw−1 (T hk−1 ) + T P Rℓw−1 (T hk )] · [F P Rℓw−1 (T hk ) − F P Rℓw−1 (T hk−1 )] {z } | Contribution of buffer ℓw−1
(26) VUS-PR Calculation. The volume under the Precisionℓ -Recallℓ -ℓ surface is: L N 1 X X (k,w) VUS-PR = ∆ · |ℓw − ℓw−1 | 2 w=1
(27)
k=1
∆(k,w) = Precisionℓw (T hk ) · [Recallℓw (T hk ) − Recallℓw (T hk−1 )] {z } | Contribution of buffer ℓw
+ Precisionℓw−1 (T hk ) · [Recallℓw−1 (T hk ) − Recallℓw−1 (T hk−1 )] {z } |
(28)
Contribution of buffer ℓw−1
B.3
Setting
In our experiments, we implemente SCAN using PyTorch, and all experiments are conducted on an NVIDIA GeForce RTX 3090 24GB GPU. The optimization is performed using the Adam optimizer with an initial learning rate of 10−3 . We set the batch size as 32, the dimension of hidden states as 256, and the multi-scale as {25, 5, 1}. We employ a sliding window of length 2500 for time series processing and perform anomaly detection based on non-overlapping windows. To ensure a fair performance comparison, we refrain from applying the drop last trick in the inference phase [70]. After obtaining the anomaly scores, we use the widely used SPOT [53] method to determine the threshold.
C
Additional Experiments and Analysis
C.1
Multi-metrics Results
For a more comprehensive comparison, we compare our SCAN model with four recognized advanced methods across multiple metrics, including Accuracy (Acc), Affiliation-F1 (Aff-F1), AUC-ROC (A-R), AUC-PR (A-P), Range-AUC-ROC (R-A-R), Range-AUC-PR (R-A-P), VUS-ROC (V-R) and VUS-PR (V-P), as shown in Table 4. SCAN achieves superior or comparable performance over all metrics, further validating its effectiveness. C.2
UCR Benchmark
The UCR dataset [57] consists of 250 sub-datasets. Each sub-dataset contains a univariate time series with only a single anomaly segment. For the UCR dataset, the time series anomaly detection task aims to identify the location of the anomaly within the test set of each sub-dataset. Therefore, we compute the anomaly score for each time step in the test set and then rank them in descending order. Following the approach of Timer [71], anomaly detection is accomplished if the time step with the alpha quantile hits the labeled anomaly interval in the test set. We evaluate SCAN on the UCR dataset, comparing it against Timer [71], DADA [10], and CrossAD [11], as shown in Figure 6. The left figure shows the number of datasets in which the model successfully detects anomalies at the 3% and 10% quantile levels. The right figure displays the quantile distribution and the average quantile across all UCR datasets. SCAN achieves a higher number of successful detections and a lower average quantile, demonstrating its superior anomaly detection capability. Furthermore, we visualize the detection results of SCAN to highlight its robust anomaly detection capability, as shown in Figure 7 to Figure 14. 16
Table 4: Multi-metrics results in the three real-world datasets. Higher values for all metrics indicate better performance. Red: the best, Blue: the 2nd best. Dataset
Method
Acc
Aff-F1
A-R
A-P
R-A-R
R-A-P
V-R
V-P
SMD
ModernTCN TimeMixer TimesNet KAN-AD SCAN
0.9074 0.9091 0.9066 0.9161 0.9202
0.8316 0.8408 0.8397 0.8297 0.8431
0.7021 0.6795 0.7603 0.7406 0.8163
0.1401 0.1215 0.1627 0.1549 0.2182
0.7754 0.7549 0.8300 0.7682 0.8736
0.1628 0.1580 0.2320 0.1605 0.3004
0.7707 0.7711 0.8420 0.7657 0.9081
0.1596 0.1390 0.2040 0.1593 0.3203
PSM
ModernTCN TimeMixer TimesNet KAN-AD SCAN
0.3016 0.2774 0.2773 0.7347 0.6298
0.7959 0.7499 0.7970 0.8534 0.8463
0.5846 0.5522 0.5755 0.6434 0.6452
0.3843 0.3447 0.3701 0.4560 0.4478
0.6550 0.6140 0.6617 0.6375 0.7413
0.4748 0.4089 0.4827 0.4514 0.5895
0.6480 0.5974 0.6344 0.6377 0.7457
0.4689 0.3807 0.4373 0.4523 0.5851
GECCO
ModernTCN TimeMixer TimesNet KAN-AD SCAN
0.9882 0.9832 0.9860 0.8333 0.9897
0.9018 0.8989 0.8906 0.7904 0.9591
0.9595 0.9620 0.9173 0.8415 0.9811
0.4325 0.3559 0.3431 0.3259 0.5491
0.9733 0.9803 0.9327 0.8642 0.9748
0.5036 0.5106 0.3593 0.1755 0.5339
0.9694 0.9899 0.9834 0.8567 0.9957
0.4819 0.4106 0.4578 0.1732 0.6812
Figure 6: Anomaly detection results based on the UCR benchmark. The left figure shows the number of datasets in which the model successfully detects anomalies at the 3% and 10% quantile levels, with higher counts indicating stronger detection capability. The right figure displays the quantile distribution of SCAN and the average quantiles of all models across all UCR dataset, with lower average quantiles indicating stronger detection capability. C.3
Complexity Analysis
Let j be the number of scales, Ni be the number of patches at the i-th scale, demb be the embedding dimension (demb = d × C), dr be the clustering space dimension, K be the number of clusters, and L be the patch size. The computational cost of SCAN is the sum of the complexities of its constituent modules across j scales and dominated by the following core parts: • Representation extraction: Constructing the similarity adjacency matrix Asim involves i computing all-pairs cosine similarities, which takes O(Ni2 demb ). The temporal-dependent 2 adjacency matrix Atim i is constructed in O(Ni ). Neighborhood aggregation and centering operations involve matrix multiplications of size (Ni × Ni ) and (Ni × demb ), resulting in a complexity of O(Ni2 demb ). • Clustering: Projecting patches into the clustering space takes O(Ni demb dr ). The computation of the soft assignment Yisoft requires calculating distances between Ni patches and K centroids, taking O(Ni Kdr ). Constructing the same-cluster probability matrix Si requires O(Ni2 K). The clustering loss Lclu involves matrix operations with Si and the membership matrix Mi , which scale as O(Ni2 K). • Fusion & Reconstruction: This module first generates cluster-weighted representations by aggregating K cluster centers according to the membership probabilities, which takes O(Ni Kdr ). The fused representations are then mapped back to the original input space via a linear head. Since each of the Ni patches is projected to a patch of length L, the mapping cost is O(Ni Ldemb ). 17
Thus, the overall time complexity is: j X
O
!
Ni2 (demb + K) + Ni (demb dr + Kdr + Ldemb )
(29)
i=1
C.4
Parameter Sensitivity Analysis
The primary hyperparameter affecting SCAN performance is the number of clusters. To analyze its impact on anomaly detection, we experimentally evaluate the performance of the model under different parameters. Effective clustering provides reliable support for anomaly detection. Table 5 shows the performance of the model at different cluster numbers k. It can be concluded that increasing the number of clusters generally improves model performance, with performance enhancement tending to become stable after a certain number of clusters is reached. Table 5: Parameter sensitivity analysis on cluster numbers k. Higher VUS-ROC (V-R) or VUS-PR (V-P) values indicate better performance. Bold: the best. Dataset
D
MSL
SMAP
PSM
SWaT
GECCO
Metric
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
V-R
V-P
k=5 k=10 k=15 k=20 k=25 k=30 k=35 k=40 k=45 k=50
0.8092 0.8232 0.8163 0.8187 0.8167 0.8225 0.8222 0.8228 0.8230 0.8222
0.3330 0.3315 0.3137 0.3194 0.3151 0.3191 0.3224 0.3207 0.3227 0.3202
0.5694 0.6114 0.6068 0.6020 0.6016 0.6014 0.6024 0.6005 0.5996 0.5991
0.1385 0.1578 0.1556 0.1519 0.1516 0.1521 0.1521 0.1514 0.1506 0.1511
0.7045 0.7176 0.7451 0.7418 0.7448 0.7449 0.7454 0.7457 0.7449 0.7452
0.5323 0.5643 0.5846 0.5821 0.5843 0.5835 0.5844 0.5851 0.5838 0.5848
0.5621 0.6811 0.6819 0.7443 0.7445 0.7421 0.7826 0.7866 0.7713 0.7891
0.2409 0.3832 0.2851 0.3881 0.4107 0.4881 0.4697 0.4986 0.4245 0.4395
0.9937 0.9958 0.9957 0.9957 0.9956 0.9955 0.9956 0.9957 0.9956 0.9957
0.6271 0.6785 0.6736 0.6795 0.6757 0.6698 0.6737 0.6799 0.6789 0.6812
Broader impacts
Time series anomaly detection plays a vital role in practical applications. Early detection of anomalies can prevent potential risks and reduce economic losses. This work addresses the trade-off dilemma between over-generalization and under-generalization of reconstruction-based methods, which helps promote the development of this field. Focusing on technical improvement, this work does not involve datasets that may contain sensitive information or privacy risks, and thus will not exert any negative impacts on society.
18