ConceptioArchivearXiv CS
arXiv CSopen access

Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
databasesdatamanagementsqlstorage
databases, sql, data management, storage

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

1

Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments Zhichen Lai, Huan Li∗ , Member, IEEE, Dalin Zhang∗ , Senior Member, IEEE, Dong Gong, Lina Yao, Senior Member, IEEE, and Christian S. Jensen, Fellow, IEEE,

arXiv:2607.23503v1 [cs.LG] 26 Jul 2026

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. This article has been accepted for publication in IEEE Transactions on Knowledge and Data Engineering. This is the author’s accepted version and has not been fully edited by IEEE.

Abstract—Imputation is a well-established data-cleaning task in the database community, and deep learning has advanced the state of the art. Internet of Things (IoT) applications generate vast amounts of Correlated Time Series (CTS) data that often contain missing values and therefore require imputation. However, current methods predominantly emphasize accuracy and often neglect the adaptability required in changing IoT environments. First, they struggle with sensor failures because CTS imputation relies on correlations among sensors; when some sensors are unavailable, the imputation of values from other sensors may be disrupted. Second, they cannot selectively impute specific sensors, leading to unnecessary processing of values known to be complete. Third, their static architectures cannot adapt to changing resource availability, resulting in consistently high computational demands. To tackle these limitations, we propose A DACTS I, an adaptive CTS imputer for changing environments. The architecture integrates a One-shot Temporal Convolutional Network with a Learned Time-Sensor Index Table to extract and decouple complex spatio-temporal features into sensor-wise embeddings, thereby enabling adaptation to varying sensor subsets. Sparse Spatial Attention extracts dynamic spatial correlations efficiently, while Correlation-Weighted Sensor Selection ensures sufficient spatial context by selecting informative sensors. Experimental results across twelve baseline methods, three adaptability scenarios, and five benchmark datasets spanning traffic, air quality, and trajectory domains show that A DACTS I achieves an average 33.1% reduction in MAE relative to the strongest baseline on each dataset, while supporting sensor-subset and resource-adaptive inference with a single trained model. Moreover, its modest memory footprint enables deployment on a range of commodity computing devices, including MCUs, making A DACTS I an adaptive CTS imputation solution for changing environments. The source code is available at https://github.com/ryanlaics/adactsi/. Z. Lai is with the College of Computer and Data Science, Fuzhou University, Fuzhou, China; the Engineering Research Center of Big Data Intelligence, Ministry of Education, Fuzhou, China; the Fujian Key Laboratory of Network Computing and Intelligent Information Processing, Fuzhou University, Fuzhou, China; and the Department of Computer Science, Aalborg University, Aalborg, Denmark. E-mail: [email protected]. D. Zhang is with the Space Information Research Institute, Hangzhou Dianzi University, Hangzhou, Zhejiang, China, and the Department of Computer Science, Aalborg University, Aalborg, Denmark. E-mail: [email protected]. C. S. Jensen is with the Department of Computer Science, Aalborg University, Aalborg, Denmark. E-mail: [email protected]. H. Li is with the College of Computer Science and Technology, Zhejiang University, Hangzhou, China, and the State Key Laboratory of Blockchain and Data Security, Zhejiang University, Hangzhou, China. E-mail: [email protected]. D. Gong is with the School of Computer Science and Engineering, The University of New South Wales, Sydney, NSW, Australia. E-mail: [email protected]. L. Yao is with the School of Computer Science and Engineering, The University of New South Wales, Sydney, NSW, Australia, and CSIRO’s Data61, Eveleigh, NSW, Australia. E-mail: [email protected]. *Corresponding authors: H. Li and D. Zhang.

Index Terms—Correlated Time Series, Data Imputation, Model Adaptability. S3 Dynamic-Resource Adaptability

S1 Sensor-Failure Adaptability

TS0 TS1 TS2 TS3 TS4 TS5

S2 Excessive-Imputation Adaptability

Fig. 1. Adaptability scenarios for CTS imputation: Sensor-Failure Adaptability (S::SFA) excludes failed sensors; Excessive-Imputation Adaptability (S::EIA) selectively processes incomplete sensors; and Dynamic-Resource Adaptability (S::DRA) adjusts the model and sensor-group size to available resources.

I. I NTRODUCTION Internet of Things (IoT) applications, including traffic management [1], [2], wearable devices [3], [4], and database monitoring systems [5], [6], rely on interconnected sensors to monitor operational parameters such as traffic speed in road networks. These sensors generate vast amounts of highly correlated time series data because of their spatio-temporal relationships; we refer to such data as Correlated Time Series (CTS). CTS data are often stored in spatio-temporal or temporal databases [7], [8] for downstream tasks, including classification [3], [9], anomaly detection [10], [11], and forecasting [2], [12]. In practice, sensor failures, network interruptions, and battery depletion introduce missing values that distort CTS patterns and degrade downstream analyses, as shown by the forecasting results in Table IX. The resulting false alerts or biased predictions [13] motivate accurate CTS imputation. Recent deep learning (DL) imputers achieve state-of-the-art performance [14]–[20], but typically assume fixed structures, capacities, and sensor partitions. In deployment, available FLOPs and memory vary, as does the workload1 . Complete values need no imputation, while failed or removed sensors should be excluded because their placeholders can corrupt correlation-based estimates for other sensors. Recent practices in elastic computing [21]–[24] motivate adaptive CTS imputation frameworks that adjust model config1 The workload is the number of input sensors because DL-based imputers typically use short, fixed-size windows (e.g., 24 time steps).

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

urations to online workloads and available resources. Predictive resource provisioning and scheduling provide complementary system-level mechanisms [24]–[26]; this paper focuses on the imputation model design and addresses the following question: Is it possible to design a DL-based imputer (i.e., imputation model) that can adapt to variations in online workload and available computing resources to perform on-demand inference, thereby improving resource utilization while maintaining imputation accuracy? The on-demand inference capability of such an imputer can support at least three desirable scenarios: S1 Sensor-Failure Adaptability (S::SFA). As depicted in the yellow sector of Fig. 1, failed sensors (i.e., TS0 and TS3–TS5) can significantly disrupt the imputation of values from other sensors because imputation relies on learning correlations among sensors. Existing DLbased imputers assume a constant input dimensionality, so when sensors fail, they often fill the unavailable sensor inputs with default or random values. These placeholder values can reduce the accuracy of imputing the remaining sensors and incur unnecessary computation. Hence, an imputer should adjust its input size and exclude failed sensors. By reducing the input dimensionality on the fly, the model avoids erroneous placeholder values and reduces computational demand. This dynamic, on-demand inference capability enables stable operation during sensor failures, maintenance activities, and changing online workloads, enhancing robustness and efficiency. We study this scenario experimentally in Section V-C. S2 Excessive-Imputation Adaptability (S::EIA). The red sector of Fig. 1 illustrates that when most sensors report complete data, current DL-based methods typically process all sensors uniformly and cannot easily restrict computation to the incomplete sensors (e.g., TS3), leading to unnecessary processing and excessive computation. This inefficiency is exacerbated in large-scale sensor settings, where only a few sensors have missing values. Inspired by dynamic resource allocation in elastic computing [21], there is a need for imputation approaches that can reduce the input workload while still leveraging informative complete sensors. Such selective processing minimizes unnecessary computation and allocates resources to the most informative signals, improving both efficiency and imputation quality. We study this scenario in Section V-D. S3 Dynamic-Resource Adaptability (S::DRA). The green sectors in Fig. 1 illustrate that the static structures and parameter capacities of current DL-based imputers impose consistently high computational demands and limit adaptation to changing online workloads and available resources. An adaptive imputer can instead match its computational requirements to the available resources in real time. When resources are constrained, it can reduce demand by simplifying its architecture (① to ②, deactivating some parameters) and partitioning the input (② to ③, splitting sensors into the groups TS0–

2

TS2 and TS3–TS5). Conversely, when more resources become available, it can activate additional parameters and increase the partition size to process more or all sensors simultaneously. This adaptability can improve resource utilization and imputation performance under online conditions. We study this scenario experimentally in Section V-E. Addressing these scenarios requires methods that go beyond static structures and parameter capacities by supporting variable sensor subsets and input dimensionalities. Workload …

Memory Constraint

Processing Power

Fig. 2. On-demand Mode for Handling Variations in Input Workload and Resource Availability.

As shown in Fig. 2, On-demand Mode uses one adaptable N -in-1 model and a shared parameter set across configurations, avoiding the design, training, deployment, and maintenance of many workload- or resource-specific models. We therefore present A DACTS I, an adaptive CTS imputer for on-demand inference. Its One-shot Temporal Convolutional Network (One-shot TCN) and Learned Time-Sensor Index Table (LTSIT) produce sensor-wise spatio-temporal representations. Sparse Spatial Attention (SSA) extracts informative correlations from variable sensor subsets, while CorrelationWeighted Sensor Selection (CW-SS) supplements incomplete sensors with informative complete ones when needed. The key contributions are summarized as follows. • We identify three practical adaptability challenges in CTS imputation under changing workloads and computing resources. • We propose A DACTS I , which combines One-shot TCN, LTSIT, SSA, and CW-SS to support variable sensor subsets, model capacities, and resource budgets. • Extensive experiments, including downstream forecasting, show that A DACTS I supports on-demand adaptability while consistently achieving higher accuracy than existing imputers. The remainder of this paper presents the preliminaries, design principles, method, experiments, related work, and conclusion. II. P RELIMINARIES We formalize the problem of Correlated Time Series (CTS) imputation and introduce key definitions that form the foundation of our adaptive framework. Definition 1 (Correlated Time Series, CTS). Correlated Time Series (CTS) are multiple time series generated by sensors

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

3

TABLE I D ESIGN P RINCIPLES (P1 TO P3) FOR A DAPTABILITY IN D IFFERENT S CENARIOS (S1 TO S3 IN S ECTION I) Design Principles

P1. Sensor Variation Handling

P2. Informative Sensor Selection

P3. Model Complexity Tuning

S1. Sensor-Failure Ada. S2. Excessive-Imputation Ada. S3. Dynamic-Resource Ada.

⋆ Adapt to available sensors ⋆ Process only sensors with missing data ⋆ Partition sensors to match memory capacity

Not required ⋆ Select sensors for spatial information Not required

Not required Not required ⋆ Adapt complexity to processing power

Existing DL-based Imputers A DACTS I (ours)

l Unsupported m Supported (One-shot TCN, LTSIT, SSA, CW-SS)

l Unsupported m Supported (SSA, CW-SS)

l Unsupported m Supported (SSA)

monitoring a shared set of processes. Consider a set of N sensors, each producing a time series. Let X ∈ RN×T denote these time series, where Xi,t is the value from the i-th sensor at time t [27]. In CTS, temporal correlations refer to the dependencies along the time axis within a single sensor’s series; for example, the measurement at time t can be influenced by measurements at previous time steps. In contrast, spatial correlations denote dependencies among measurements from different sensors, observed at the same time or even across time, because these sensors monitor shared or related processes2 . Following prior work (e.g., [27]), we do not impose a fixed correlation threshold; correlations can be positive or negative and are dynamic and applicationdependent in practice. Definition 2 (CTS Imputation). Given a CTS X ∈ RN×T with missing values indicated by a mask matrix M ∈ {0, 1}N×T (where ”0” denotes missing), CTS imputation estimates the missing values to improve data quality. The imputation process replaces only missing values. X̂ = X ⊙ M + g(X | M) ⊙ (1 − M),

(1)

where X̂ ∈ RN×T is the imputed CTS, ⊙ denotes elementwise multiplication, 1 ∈ RN×T represents a matrix of ones, and g(X | M) is the imputation function that estimates the missing values based on the observed data X and the mask M. Definition 3 (Adaptive CTS Imputation). Adaptive CTS imputation refers to a model’s ability to adjust its architecture and input processing in response to changing workloads and resource availability to ensure accurate, ondemand imputation. The required forms of adaptability include Sensor-Failure Adaptability, Excessive-Imputation Adaptability, and Dynamic-Resource Adaptability. Specifically, an imputer should be able to exclude failed sensors, selectively process incomplete subsets of sensors, and adapt its resource requirements by adjusting both its architecture and the size of input sensor partitions on demand. The adaptive imputation process is defined as follows. X̂cS = XS ⊙ MS + g c (XS | MS ) ⊙ (1 − MS ) ,

(2)

where S ⊆ {1, 2, . . . , N} represents the subset of sensors being processed and XS ∈ R|S|×T contains the data from this subset. MS ∈ {0, 1}|S|×T is the corresponding mask matrix. The imputation function g c (XS | MS ) estimates the missing 2 Cross-sensor correlations can reflect more than spatial proximity (e.g., functional or process-based dependencies), but following convention, we refer to them as spatial correlations [28].

values, with model complexity level c determining its resource demands. The matrix 1 has the same shape as MS . Finally, X̂cS denotes the imputed CTS for the selected sensors under input size |S| and complexity level c.

Adaptive Imputation Accuracy Gain

Non-adaptive Imputation

T1

T2

T3

T1

T2

Input Workload

Imputation Accuracy

Model Complexity

Resource Demand

T3

Fig. 3. Non-adaptive vs. Adaptive Imputation under Stable Resource Demand.

Fig. 3 illustrates how the imputer adjusts its complexity to the input workload to maintain stable resource demand. In this scenario, the left diagram represents non-adaptive imputation, where both the input workload and model complexity remain constant at time points T 1, T 2, and T 3, resulting in a relatively stable accuracy level. In contrast, the right diagram shows adaptive imputation, in which the imputer dynamically increases its complexity in response to lower workloads at T 1 and T 3, thereby improving accuracy under a fixed resource budget. The shaded areas in Fig. 3 highlight this accuracy gain. III. D ESIGN P RINCIPLES As discussed in Section I, CTS imputation in realworld applications encounters three primary challenges: Sensor-Failure Adaptability (S::SFA), Excessive-Imputation Adaptability (S::EIA), and Dynamic-Resource Adaptability (S::DRA). While existing DL-based approaches have achieved strong CTS imputation performance, they assume a fixed model structure and parameter capacity during deployment. This rigidity causes inefficiency and performance degradation when sensor availability, imputation demand, or computing resources change. To broaden the applicability of imputation, we propose three key design principles, each addressing a specific challenge and guiding adaptation to changing conditions. Together, these principles form the foundation of A DACTS I for adaptive CTS imputation.

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

P1 Sensor Variation Handling (P::SVH). In IoT applications, the available sensor set may vary because of failures, maintenance activities, or changing online workloads. Moreover, in many practical scenarios, only a small subset of sensors contains missing data, while the remaining sensors provide complete data. Existing DL-based imputers nevertheless process all sensors uniformly, including failed or unnecessary ones, which can degrade imputation performance and increase computational overhead. P::SVH addresses this limitation by selectively processing relevant sensors. To achieve such adaptability, an imputer should decouple complex spatio-temporal features into sensor-wise representations, enabling effective imputation under sensor variation. P2 Informative Sensor Selection (P::ISS). Traditional DLbased imputers process data from all sensors regardless of their characteristics, leading to excessive and inefficient computation. P::ISS mitigates this problem by identifying the most informative sensors and using their data for spatial feature extraction. To achieve P::ISS, an imputer should adopt metrics such as sensor-wise similarity to evaluate each sensor’s informativeness. Focusing on key sensors can reduce computation while preserving imputation accuracy. P3 Model Complexity Tuning (P::MCT). Deploying imputation models across diverse environments requires models with adjustable computational demands. Existing DL-based imputers maintain a constant model structure and parameter capacity, resulting in consistently high computational demands. To achieve P::MCT, an imputer should support on-demand complexity tuning based on available resources. By dynamically activating subsets of model parameters or selectively processing feature maps, an imputer can scale its computational load and operate across varying resource levels. As illustrated in Table I, an imputer that adopts the three design principles can adapt to sensor variation, prioritize informative sensors for effective imputation, and adjust computational demand according to available resources. IV. A DACTS I D ESIGN Building on the previously discussed design principles, we introduce A DACTS I, an adaptive data imputer for CTS. A DACTS I is specifically designed to address Sensor-Failure Adaptability (S::SFA), Excessive-Imputation Adaptability (S::EIA), and Dynamic-Resource Adaptability (S::DRA). A. The Overall Architecture of A DACTS I As depicted in Fig. 4(a), A DACTS I comprises several key components that jointly enable on-demand imputation across dynamic scenarios: • The input data are first processed by the CorrelationWeighted Sensor Selection (CW-SS) mechanism, which selects the most spatially informative sensors for on-demand imputation when the set of incomplete sensors is too small to provide sufficient spatial correlation information (P::ISS). • Next, the One-shot Temporal Convolutional Network (One-shot TCN) processes the selected sensor data, as

4

shown in Fig. 4(b). This module extracts temporal features and employs a one-shot mechanism to reduce temporal dimensionality. It decouples complex spatio-temporal features into one-dimensional sensor-wise representations (P::SVH), allowing subsequent spatial modules to operate without managing temporal variations. • Subsequently, the Learned Time-Sensor Index Table (LTSIT) module (Fig. 4(c)) retrieves learned embeddings through multi-view indexing. This module captures periodic temporal patterns, sensor-specific characteristics, and local intra-window features via table lookups, thereby constructing comprehensive sensor-wise latent representations (P::SVH). • Finally, as demonstrated in Fig. 4(d), the Sparse Spatial Attention (SSA) module captures spatial correlations without channel mixing, preserving distinct sensor features and preventing sensor-wise feature fusion. The SSA adjusts its attention sparsity rate to enable P::MCT, focusing on data from the most spatially informative sensors for more effective spatial correlation extraction (P::ISS). The output module then generates the final imputed data. Together, these four components enable adaptive CTS imputation. We describe CW-SS last because it depends on the embeddings generated by LTSIT. B. One-shot Temporal Convolutional Network As depicted in Fig. 4(b), the One-shot TCN extracts and compresses temporal features for each sensor by transforming input feature maps of size N × T × D (number of sensors × window length × embedding size) into a compact N × D representation. This compression decouples complex spatiotemporal features into sensor-wise one-dimensional vector representations and simplifies the design of subsequent spatial modules, which no longer need to account for temporal variations, thereby supporting the P::SVH principle. The One-shot TCN employs dilated causal convolutions [29] to capture both short- and long-term temporal patterns within the input window. The input, denoted as EI ∈ RN×T×(2+dw ) , comprises the concatenation of the input CTS X ∈ RN×T , the mask matrix M ∈ RN×T , and the intra-window embedding EW ∈ RN×T×dw from the LTSIT (discussed in Section IV-C). This input is first embedded into a latent representation H ∈ RN×T×dd using a CNN embedding layer:  H[i, t, d] = ReLU EI [i, t, :]Wd + bd , (3) where Wd ∈ Rke ×(2+dw )×dd is the convolutional filter, bd ∈ Rdd is the bias term, and ke is the kernel size. Next, the TCN is applied: TCN(H | δ, kt ) = H′ , H′ [i, t, d] =

kX t −1

where  H[i, t − δ × p, :] · W d [p, :] ,

p=0

where H′ ∈ RN×T×dd is the latent representation of the TCN, W d [p, :] ∈ Rdd is the d-th convolutional filter at position p, kt is the kernel size, and δ is the dilation rate.

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

TS0 NULL NULL NULL NULL NULL

TS1

2.3

2.5

3.5

3.9

TS2

4.0

0.2

0.5

0.1

0.9

0.2

TS3 NULL

3.5

3.3

3.0

NULL

TS4

1.0

1.0

1.1

1.3

1.2

TS5

8.1

1.3

3.5

9.9

1.2

TS6

0.1 T0

NULL NULL NULL T1 T2 T3

0.1 T4

5

TS0 NULL NULL NULL NULL NULL

CW-SS Selection (Sec. 4.6)

TS3 NULL

3.5

3.3

3.0

NULL

TS5 8.1

1.3

3.5

9.9

1.2

TS6 0.1

NULL NULL NULL

0.1

T1

T4

T0

T2

T3

Timestamp/Sensor Index

Sparse Spatial Attention (Sec. 4.4)

One-shot TCN (Sec. 4.2)

Retrieve Learned Time-Sensor Index Table (Sec. 4.3)

Query

Retrieve embeddings from 𝐈𝐓𝑆

TS0 1.2

1.3

2.5

2.7

3.0

TS3 2.9

3.5

3.3

3.0

2.8

TS5 8.1

1.3

3.5

9.9

1.2

TS6 0.1

0.2

0.3

0.2

0.1

T0

T1

T2

T3

T4

(a) Overall Architecture Discarded features

Index (ToD/Dow)

Embedding

Index (Sensor)

Embedding

0 (0:00)

0 (TS0)

1 (1:00)

1 (TS1)

2 (2:00)

… Zero padding

t+1

Input window

t+T

(b) One-shot TCN

Output Module (Sec. 4.5)

Projection

𝐐

𝐊

𝐕

Spatial Informativeness Measuring

𝐈𝐓𝑆

24 (Mon)

Step Sensor

0

1

2

Sensor Filtering

𝐐𝐫𝐞𝐝𝐮𝐜𝐞𝐝

25 (Tue)

0 (TS0)

26(Wed)

1 (TS1)

𝐒𝐫𝐞𝐝𝐮𝐜𝐞𝐝

Feedforward Network

Sparse Attention Calculation

𝐈𝐓 𝑊

𝐈𝐓 𝑇

(c) Learned Time-Sensor Index Table

(d) Sparse Spatial Attention

Fig. 4. (a) Overview of A DACTS I. CW-SS optionally selects input sensors when the set of incomplete sensors is too small to provide adequate correlation information. (b) One-shot TCN then extracts and compresses temporal features, which are enriched with multi-view spatio-temporal embeddings from (c) LTSIT. Next, (d) the SSA module applies self-attention to the most informative sensors, capturing dynamic spatial correlations while improving computational efficiency. Finally, the output module independently generates imputed values for each sensor.

A one-shot mechanism is then applied, retaining only the feature at the final time step: ED [i, d] = H′ [i, −1, d].

of Day (ToD) from the input window’s first timestamp, is mapped to a corresponding embedding vT ∈ Rdt . This embedding is broadcast across all sensors. Formally, the temporal embeddings are represented as:

(4)

The resulting compressed representation ED ∈ RN×dd encapsulates each sensor’s temporal features in a more manageable and compact form. C. Learned Time-Sensor Index Table Inspired by decoupled learning [18], [30], we propose a Learned Time-Sensor Index Table (LTSIT) to extract and decouple spatio-temporal patterns from multiple views, as shown in Fig. 4(c). The LTSIT supports P::SVH by enabling on-demand embedding retrieval with constant time complexity (O(1)) via table lookups. These embeddings are learned from the training data and indexed by sensor, timestamp, or inputwindow position. Once trained, the LTSIT directly retrieves multi-view feature representations during inference without additional feature-extraction computation. This design makes it highly adaptable to sensor variation (P::SVH). Unlike prior work that typically handles either temporal or spatial patterns, LTSIT’s key innovation lies in its joint learning of periodic temporal, sensor-specific, and intra-window patterns through three specialized index tables. T • Periodic Temporal Index Table (IT ): Captures periodic temporal patterns. Each unique temporal key kT , formed by periodic indices such as Day of Week (DoW) and Time

ITT : {kT → vT | kT = [DoW, ToD, . . .], vT ∈ Rdt }, (5) ET = [vT ]×N , ET ∈ RN×dt . (6) Here, ET represents the temporal embedding, with each sensor sharing the same embedding vT derived from ITT . S • Sensor Identity Index Table (IT ): Encodes sensorspecific patterns by assigning each sensor identity to a unique embedding vnS ∈ Rds , capturing inherent characteristics of different sensors: ITS : {kS → vnS | kS ∈ {1, . . . , N},

vnS ∈ Rds },

(7)

resulting in the spatial embedding: ES = [[vnS ]⊤ | n = 1 → N], •

ES ∈ RN×ds .

(8)

W

Spatio-Temporal Window Index Table (IT ): Captures local intra-window interactions by assigning a positionW aware embedding vn,t ∈ Rdw to each data point Xn,t in the input window. This allows A DACTS I to learn patterns across sensors and time steps within the window. Formally, the window embeddings are represented as: W ITW : {kW → vn,t | kW = (n = 1 → N, t = 1 → T), W vn,t ∈ Rdw },

(9)

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

6

O(N 2D)

O(NUD) …

O(N 2 D)

Features of all sensors

N×D

×

×

D×N

N×D

=

Softmax Q

KT

× N×D

N×N

Update features O(N 2D + N 2D) = O(N 2 D)

V

(a) Vanilla Self-attention

Features of the most spatially informative sensors

O(NUD)

×

D×N

×

N×d

× N×D

=

U×D

U×N

Sensor filtering Softmax

Qreduced

KT

Update features O(NUD + NUD) = O(NUD)

V

(b) Sparse Spatial Attention

Fig. 5. Computational Processes of (a) Vanilla Self-Attention [31] and (b) Sparse Spatial Attention. Here, U is the number of sensors selected from the N total sensors and represents a sparsity level tied to processing power, D is the embedding dimension, and O(·) denotes time complexity.

leading to the window embedding: W EW = [[vn,t | t = 1 → T] | n = 1 → N], EW ∈ RN×T×dw . (10)

The window embedding EW is concatenated with the input CTS X and the mask matrix M to form the input for the One-shot TCN. Subsequently, the embeddings ET and ES are concatenated with the output of the One-shot TCN, ED , to create a comprehensive feature representation: D

S

T

E = Concat(E , E , E ),

E∈R

N×D

,

(11)

where D = dd + ds + dt . The concatenated representation E combines compressed temporal features with periodic temporal, sensor-specific, and intra-window embeddings, delivering a streamlined sensorwise feature representation to subsequent modules. Specifically, LTSIT’s constant-time retrieval mechanism facilitates sensor-agnostic feature extraction. This approach also supports the P::SVH principle by effectively addressing sensor variation. D. Sparse Spatial Attention While some existing studies adopt self-attention for spatial correlation extraction [16], [18], [19], they assume a constant input sensor dimensionality, limiting their adaptability to scenarios with sensor variation (P::SVH). Even if failedsensor inputs are filled with zeros, these methods suffer from degraded imputation accuracy and incur unnecessary computational costs, as evidenced by Table VI. To address this limitation, we introduce the Sparse Spatial Attention (SSA) module, designed to capture dynamic spatial correlations among sensors under varying input conditions. As shown in Fig. 5, unlike traditional self-attention mechanisms that require the complete sensor set, SSA performs adaptive sparse self-attention solely on the sensor subset, without channel mixing. This design avoids sensor-wise feature fusion, enabling the module to adapt to sensor variation in line with the P::SVH principle. By employing a sparse attention mechanism, SSA identifies and prioritizes the most spatially informative sensors, enhancing both imputation efficiency and accuracy (P::ISS). Moreover, SSA can adjust its sparsity level based on available processing power, enabling the imputer to scale its computational demands and operate across devices with different processing capacities (P::MCT). The SSA module follows these steps: 1) Self-Attention Projection: Computes the queries Q, keys K, and values V from the combined feature E ∈ RN×D : Q = EWQ ,

K = EWK ,

V = EWV ,

(12)

where WQ , WK , WV ∈ RD×D are learned projection matrices. 2) Spatial Informativeness Measurement: Measures the spatial informativeness of each sensor to identify the most informative ones (P::ISS). For a query qi and keys K, the measure M(qi , K) is defined as: M(qi , K) = ln

N X j=1

e

T qi kj √ D

N

1 X qi kjT √ , N j=1 D

(13)

where the first term measures the compatibility between query qi and all keys kj . The exponential emphasizes strongly aligned query-key pairs, while the logarithm controls the scale and prevents large values from dominating. The second term subtracts the mean interaction, smoothing the measure and enabling more balanced comparisons among sensors. 3) Sensor Filtering: During training, a random sparsity scheme selects a subset of U sensors from the N sensors, where U ≤ N is a randomly chosen integer corresponding to the sparsity rate. During inference, U is set according to the computational requirements, and the top U sensors according to M(qi , K) are selected for self-attention. 4) Sparse Attention Calculation: Performs self-attention on the filtered sensors to compute attention scores and derive output features A:  reduced ⊤  Q K reduced √ A(Q , K, V) = Softmax V, (14) D where Qreduced ∈ RU×D denotes the queries of filtered sensors and K, V ∈ RN×D are the keys and values, respectively. 5) Add & Norm and Feed-Forward Network: Applies residual connections and layer normalization: Z = LayerNorm(A + E).

(15)

The module then passes the result through the feed-forward network for nonlinear transformation: FFN(Z) = LayerNorm(ReLU(ZW1 + b1 )W2 + b2 + Z), (16) where W1 , W2 ∈ RD×D are weight matrices and b1 , b2 ∈ RD are bias terms. The SSA module enhances adaptability in CTS imputation by accommodating varying numbers of input sensors (P::SVH). By using sparse spatial attention to focus on the most informative sensors (P::ISS), SSA robustly extracts spatial correlations even from a reduced sensor set. In addition, the random

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

sparsity training scheme allows SSA to maintain efficiency and accuracy across diverse input workloads (P::SVH and P::MCT), addressing the limitations of existing spatial feature extraction modules under sensor variation [2], [16], [19], [28].

7

2) Sensor Selection Strategy: To ensure that the imputer receives at least K sensors for adequate spatial information, we implement the following strategies: 1) Determine the Number of Additional Sensors Needed: Nneeded = K − Nmissing ,

E. Output and Training of A DACTS I In the final step, a multilayer perceptron (MLP) transforms the N × D features produced by SSA into an imputed output of shape N × T. Each sensor’s features are processed independently. X̂ = MLP(XSSA ),

(17)

where XSSA ∈ RN×D is the output from the SSA module, and X̂ ∈ RN×T is the final imputed output. During training, the SSA module uses a randomly selected number of sensors for sparse spatial attention, ranging from 1 to N (P::SVH). A DACTS I uses the Masked Mean Absolute Error (Masked MAE) [16], [19], [32] as the loss function. This metric focuses on errors in the masked (missing) regions to optimize the model: PN Masked MAE =

i=1

PT

j=1 (1 − Mi,j ) ·

PN

i=1

PT

Xi,j − X̂i,j

j=1 (1 − Mi,j )

,

(18) where Xi,j and X̂i,j represent the ground-truth and imputed values, respectively, at the i-th sensor and the j-th timestamp.

F. Correlation-Weighted Sensor Selection To balance efficiency with the need for sufficient spatial correlation information, we introduce the Correlation-Weighted Sensor Selection (CW-SS) mechanism. When only a few sensors have missing values, those sensors alone may provide insufficient spatial context for accurate imputation. Conversely, using all available sensors can lead to excessive computational overhead. CW-SS addresses this challenge by ensuring that a minimum number of sensors, denoted by K, are used for ondemand imputation. The parameter study of K is presented in Section V-J, where we provide recommendations for selecting K. If the number of incomplete sensors falls below K, CWSS selectively incorporates additional sensors based on their correlations with the incomplete ones (P::ISS). As a result, CW-SS keeps imputation efficient and accurate even with a limited input sensor set (P::SVH). 1) Correlation Calculation: The correlations between sensors are determined using the Sensor Identity Index Table ITS described in Section IV-C. In this index table, each row corresponds to the embedding of a specific sensor. The correlation matrix S is calculated as follows. S = ITS × (ITS )⊤ ,

(19)

where the element Si,j represents the correlation between sensors i and j. This matrix helps identify sensors that provide significant spatial information about the input sensors.

(20)

where Nmissing is the number of incomplete sensors. 2) Compute Sensor Correlation: For each complete sensor, compute its average correlation with the incomplete sensors: X 1 (21) cosine similarity(Ej , Ei ), Sj = Nmissing i where Ej is the embedding of complete sensor j, and Ei is the embedding of incomplete sensor i. 3) Rank and Select Sensors: Rank the complete sensors by their average correlation scores Sj and select the top Nneeded sensors. The effectiveness of CW-SS is supported by several considerations. By using the learned embeddings from the Sensor Identity Index Table to calculate correlations among sensors, CW-SS selects informative sensors that are spatially aligned with those containing missing data, thereby maintaining high imputation accuracy and efficiency [33]. Ensuring that at least K sensors are used addresses insufficient spatial context, as theoretical work on data imputation emphasizes the importance of using enough correlated sensors for accurate reconstruction [33]. Our experimental results in Fig. 6 demonstrate that incorporating spatial information from more sensors improves imputation accuracy. Although random sensor selection is simpler, it may include sensors that are uncorrelated with the target sensors or otherwise misleading, thereby degrading imputation accuracy [34]. As shown in Table VII, this random strategy performs worse than CW-SS. By enforcing a minimum sensor threshold and selecting only the most spatially informative sensors, CW-SS uses spatial information efficiently. This makes A DACTS I wellsuited for applications where missing rates are relatively low or missing values are concentrated within a few sensors. G. Inference Complexity Analysis We analyze the per-window inference complexity. Let N be the number of sensors, T the window length, D the embedding size, and U the number of sensors selected by SSA (U ≤ N). The One-shot TCN with L temporal convolution layers and kernel size kt costs O(NTDLkt ) time and O(NTD) memory for intermediate features. LTSIT retrieval is constant-time per sensor, yielding O(ND) time and O((|ITS |+|ITT |+|ITW |)D) memory for the tables. SSA computes projections in O(ND2 ) and sparse attention in O(UND), reducing the standard O(N2 D) cost to linear in U. The output MLP applied per sensor costs O(NDT). CW-SS computes correlations between incomplete and candidate sensors in O(ND) and selects the top sensors in O(N log N) (or O(N) with partial selection). Overall, inference runtime and memory scale linearly with U and the effective input size; when inputs are partitioned into g groups, the per-pass inference cost scales with N/g, enabling on-demand

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

8

TABLE II OVERVIEW OF BASELINE M ETHODS . Category

Method

Description

Traditional

MEAN MICE [35] MF TRMF [36] SMV-NMF [37]

Imputes missing values in a data chunk using the mean of each dimension. Imputes missing values using multiple imputation by chained equations. Treats a data chunk as a matrix and uses iterative singular value decomposition to impute missing values. Factorizes high-dimensional time series with temporal autoregressive regularization on latent factors. Applies spatial multi-view NMF to impute missing values.

Generative

R GAIN [38] P O G E VON [16]

Uses a GAN with bidirectional encoder and decoder networks to generate missing values. Employs a VAE with a decoder integrating GRU, self-attention, and MPNNs to impute missing values.

Time-Series-Based

BRITS [32] SAITS [19] T IMES N ET [39] GRIN [40]

Uses bidirectional RNNs with masking to iteratively impute missing values. Leverages self-attention to extract spatio-temporal patterns for CTS imputation. Transforms 1D time series into 2D tensors to capture multi-period temporal variations using CNNs. Integrates MPNNs for spatial pattern identification and RNNs for temporal patterns.

Graph-based

NET3 [41]

Employs GCNs for spatial dependencies and RNNs for temporal dependencies.

complexity tuning under resource constraints. Training follows the same asymptotic order with a constant-factor overhead due to backpropagation. V. E XPERIMENTS This section presents the experimental evaluation of the proposed method. Section V-A describes the experimental setup, Section V-B compares overall performance, and Sections V-C to V-E empirically validate S::SFA, S::EIA, and S::DRA, respectively. Section V-F examines downstream forecasting performance; Sections V-G and V-H study missing-rate robustness and efficiency, respectively; Section V-I presents ablation studies; and Section V-J analyzes key parameters and provides selection recommendations. Collectively, these experiments validate the accuracy, adaptability, robustness, and efficiency of A DACTS I. A. Experimental Setup Our implementation uses PyTorch, with model architectures, datasets, and training protocols available in our opensource codebase [42]. Experiments were conducted on a highperformance computing system with an NVIDIA Quadro RTX 8000 GPU and an Intel Xeon Gold 5215 CPU at 2.50 GHz. Unless otherwise specified, we use the Adam optimizer (learning rate 1×10−3 , weight decay 0), a batch size of 32, gradient clipping at 5, cosine annealing with ηmin = 1 × 10−4 , and early stopping with patience 80. We train for 400 epochs on PeMS and 300 epochs on Beijing-AQI and Vessel. All reported results are averaged over five runs. Datasets. We use five real-world CTS datasets spanning traffic, air-quality monitoring, and vessel trajectory domains. The traffic datasets are collected from the Caltrans Performance Measurement System [43]. PeMS-BA, PeMS-LA, and PeMSSD correspond to the Bay Area, Los Angeles, and San Diego regional subnetworks, respectively, with records sampled every 5 minutes. The raw regional files contain 1,632, 2,383, and 674 candidate sensors; following prior CTS imputation protocols [16], [18], each experiment uses 64 selected sensors and applies a 25% point-missing evaluation mask unless otherwise specified. Beijing-AQI contains hourly PM2.5 readings from 12 monitoring stations in Beijing from 2013 to 2017 [44]. Unlike

TABLE III K EY S TATISTICAL P ROPERTIES OF DATASETS . Dataset PeMS-BA PeMS-LA PeMS-SD Beijing-AQI Vessel

Nodes 64 64 64 12 20

Mean±Std

CV

|ρ| mean/p90

Density

Avg deg.

Clust.

394.991±225.045 517.405±273.159 374.854±239.430 79.793±80.822 62.279±0.314

0.570 0.528 0.639 1.013 0.005

0.820/0.922 0.840/0.940 0.871/0.971 0.886/0.957 0.830/0.993

0.078 0.065 0.114 0.470 0.916

4.906 4.094 7.156 5.167 17.400

0.073 0.119 0.031 0.646 0.986

PeMS, its evaluation mask is generated by temporal inference, where values observed in the current year but missing in an adjacent year are used as evaluation targets. This yields an evaluation missing rate of approximately 2%. Vessel contains AIS measurements from 20 vessels sampled every 10 seconds [45]. It has a natural missing rate of 17.49%, and we additionally apply a sparse random evaluation mask over approximately 0.4% of the observed points. Dataset Characterization. Table III reports complementary statistical and topological properties that shape imputation difficulty. Graph statistics use the provided adjacency when available and otherwise a correlation graph. Within PeMS, the table highlights differences in traffic scale, variability, correlation distribution, and topology: PeMS-LA has the highest mean flow, PeMS-SD has the largest coefficient of variation and densest selected-sensor graph, and PeMS-LA and PeMS-BA have sparser selected-sensor graphs than PeMSSD. Beijing-AQI combines high relative variability with dense station correlations, whereas Vessel has extremely low coefficient of variation and high local smoothness. These contrasts provide a cross-domain evaluation setting where imputers must handle different value scales, correlation strengths, and graph structures. Metrics. We evaluate accuracy using Mean Absolute Error (MAE), Mean Squared Error (MSE), and Mean Relative Error (MRE). MAE measures overall accuracy, MSE emphasizes larger deviations, and MRE assesses relative accuracy, especially when ground-truth values are low. Lower values indicate better imputation performance. Baselines. We include twelve methods for overall comparison, chosen to represent a broad range of approaches. Specifically, MEAN, MF, MICE [35], TRMF [36], and SMV-NMF [37] exemplify classical non-DL approaches, including temporally regularized and spatial multi-view matrix factorization.

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

9

TABLE IV OVERALL ACCURACY C OMPARISON (PART I). PeMS-BA

Model

PeMS-LA

PeMS-SD

MAE

MSE

MRE

MAE

MSE

MRE

MAE

MSE

MRE

MEAN MF MICE TRMF SMV-NMF

192.047 ± 0.000 57.265 ± 1.148 50.861 ± 0.765 71.259 ± 0.331 48.347 ± 0.550

47504.159 ± 0.000 8091.407 ± 185.123 6724.148 ± 109.829 10281.163 ± 56.114 5874.016 ± 143.257

0.474 ± 0.000 0.141 ± 0.003 0.126 ± 0.002 0.176 ± 0.001 0.120 ± 0.001

216.681 ± 0.000 77.339 ± 0.699 64.018 ± 1.015 79.180 ± 0.301 59.412 ± 0.550

62664.657 ± 0.000 15202.678 ± 156.348 10822.355 ± 405.410 13569.843 ± 60.300 9647.162 ± 281.867

0.406 ± 0.000 0.145 ± 0.001 0.120 ± 0.002 0.149 ± 0.001 0.112 ± 0.001

208.192 ± 0.000 45.811 ± 0.318 38.978 ± 1.036 61.226 ± 0.252 38.378 ± 0.591

55780.002 ± 0.000 6044.345 ± 72.976 4771.186 ± 92.335 8316.444 ± 36.472 4616.382 ± 72.088

0.529 ± 0.000 0.117 ± 0.001 0.100 ± 0.003 0.156 ± 0.001 0.098 ± 0.002

BRITS R GAIN SAITS T IMES N ET GRIN NET3 P O G E VON

30.274 ± 0.095 38.862 ± 0.752 46.567 ± 0.530 25.859 ± 0.115 30.057 ± 1.073 35.671 ± 0.111 22.194 ± 0.046

2942.411 ± 16.511 3422.914 ± 61.281 5412.574 ± 161.132 1676.843 ± 16.144 1922.072 ± 74.327 2735.574 ± 6.138 1248.681 ± 4.297

0.075 ± 0.000 0.096 ± 0.002 0.115 ± 0.001 0.064 ± 0.000 0.074 ± 0.003 0.089 ± 0.000 0.055 ± 0.000

36.921 ± 0.133 49.611 ± 1.083 61.896 ± 0.892 27.452 ± 0.114 47.835 ± 2.059 37.652 ± 0.113 23.905 ± 0.245

3681.595 ± 21.635 5533.964 ± 234.335 10998.854 ± 204.345 2058.227 ± 6.213 4561.512 ± 298.533 3416.784 ± 6.765 1714.962 ± 31.035

0.069 ± 0.000 0.093 ± 0.002 0.116 ± 0.002 0.052 ± 0.000 0.090 ± 0.004 0.071 ± 0.000 0.045 ± 0.000

21.232 ± 0.059 33.212 ± 1.475 34.117 ± 0.886 21.583 ± 0.085 41.001 ± 1.543 34.111 ± 0.184 18.990 ± 0.112

1563.234 ± 28.309 2341.466 ± 98.314 4101.397 ± 152.141 1284.300 ± 21.839 3000.012 ± 201.018 2487.581 ± 9.798 951.559 ± 8.264

0.054 ± 0.000 0.085 ± 0.004 0.087 ± 0.002 0.055 ± 0.000 0.105 ± 0.004 0.087 ± 0.000 0.048 ± 0.000

A DACTS I

19.433 ± 0.062

1056.990 ± 5.843

0.048 ± 0.000

22.141 ± 0.178

1654.584 ± 18.762

0.041 ± 0.000

15.945 ± 0.095

754.236 ± 6.127

0.041 ± 0.000

R GAIN [38] and P O G E VON [16] use generative techniques such as GANs and VAEs. Models such as BRITS [32], T IMES N ET [39], GRIN [40], NET3 [41], and SAITS [19] integrate layered temporal and spatial feature extraction modules for imputation. Table II provides detailed descriptions of the baselines. Together, these baselines cover classical statistical and matrix-factorization methods as well as representative recurrent, convolutional, attention-based, graph-based, and generative imputers. We evaluate them as fixed-input imputers under their standard protocols.

B. Overall Performance Comparisons To comprehensively evaluate the performance of A DACTS I, we conduct extensive comparisons against representative methods on five benchmark datasets spanning traffic, air quality, and trajectory domains. TABLE V OVERALL ACCURACY C OMPARISON (PART II). Beijing-AQI

Model

Vessel

MAE

MSE

MRE

MAE

MSE

MRE

MEAN MF MICE TRMF SMV-NMF

63.953 ± 0.000 22.897 ± 0.324 29.779 ± 1.022 60.726 ± 0.015 14.691 ± 0.838

7670.420 ± 0.000 1410.897 ± 77.050 1997.575 ± 233.637 6833.239 ± 3.863 622.225 ± 89.013

0.813 ± 0.000 0.291 ± 0.004 0.378 ± 0.013 0.788 ± 0.000 0.191 ± 0.011

0.129 ± 0.000 0.027 ± 0.001 6.391 ± 0.291 0.161 ± 0.003 0.026 ± 0.002

0.025 ± 0.000 0.001 ± 0.000 95.776 ± 9.147 0.035 ± 0.001 0.001 ± 0.000

0.002 ± 0.000 0.0004 ± 0.000 0.103 ± 0.005 0.003 ± 0.000 0.0004 ± 0.000

BRITS R GAIN SAITS T IMES N ET GRIN NET3 P O G E VON

14.157 ± 0.128 68.960 ± 2.145 72.418 ± 1.872 28.144 ± 0.452 61.435 ± 1.055 61.064 ± 0.981 25.765 ± 0.312

545.447 ± 12.631 12230.704 ± 412.550 13540.215 ± 380.440 2516.321 ± 52.180 7664.824 ± 185.420 7783.898 ± 160.330 2295.291 ± 45.882

0.180 ± 0.002 0.876 ± 0.015 0.912 ± 0.012 0.354 ± 0.006 0.781 ± 0.009 0.776 ± 0.008 0.327 ± 0.005

0.079 ± 0.003 0.203 ± 0.012 0.221 ± 0.015 0.125 ± 0.006 0.137 ± 0.008 0.119 ± 0.004 0.139 ± 0.007

0.010 ± 0.001 0.062 ± 0.005 0.075 ± 0.006 0.028 ± 0.002 0.032 ± 0.003 0.025 ± 0.002 0.035 ± 0.002

0.001 ± 0.000 0.003 ± 0.000 0.004 ± 0.000 0.002 ± 0.000 0.002 ± 0.000 0.002 ± 0.000 0.002 ± 0.000

A DACTS I

6.159 ± 0.082

95.818 ± 2.451

0.080 ± 0.001

0.007 ± 0.000

0.0001 ± 0.000

0.0001 ± 0.000

Tables IV and V demonstrate that A DACTS I consistently outperforms all baselines across the five datasets, including the matrix-factorization baselines TRMF and SMV-NMF. On the PeMS traffic datasets, which exhibit strong spatio-temporal correlations, A DACTS I achieves an average MAE improvement of approximately 12.0% over P O G E VON (e.g., 15.945 vs. 18.990 on PeMS-SD). The Beijing-AQI and Vessel datasets further illustrate how different value scales, correlation structures, and missing patterns affect imputation accuracy. On Beijing-AQI, BRITS and SMV-NMF are the strongest non-A DACTS I baselines in the deep and classical groups, respectively. Several deep models show large MRE values under temporal inference, indicating limited cross-year transfer in this setting. A DACTS I achieves the best results (MAE

= 6.159, a 56.5% reduction relative to BRITS) through its decoupled LTSIT design: the Sensor Identity Index Table (ITS ) captures stable station-specific characteristics, while the Periodic Temporal Index Table (ITT ) independently models recurring temporal patterns. The Vessel dataset has very low relative variation, strong lag-1 smoothness, short sequences, and an extremely sparse 0.4% evaluation mask. These properties make simple interpolation and low-rank reconstruction well matched to the data, as reflected by the strong results of SMV-NMF and MF (MAE = 0.026 ± 0.002 and 0.027 ± 0.001, respectively). A DACTS I achieves a further improvement (MAE = 0.007) through dedicated per-vessel embeddings in ITS and SSA’s adaptive filtering of less informative cross-series interactions, combining sensor-specific modeling with deep temporal representation learning. These results validate that A DACTS I’s decoupled sensorspecific embedding design provides essential flexibility across diverse spatial correlation structures and missing patterns. C. Empirical Verification for S::SFA In real-world sensor networks, sensor failures leave some nodes unobserved and reduce the spatial context available for imputation. To assess sensor-failure adaptability (S::SFA), we conduct a PeMS-SD stress test by varying the failed-sensor ratio from 12.5% to 50%. All methods are evaluated under the same failed-sensor sets. We compare A DACTS I with three representative CTS imputers: P O G E VON, T IMES N ET, and BRITS, covering generative/self-attention, convolutional, and recurrent designs. For the fixed-input baselines, failed sensors are represented by zeroed values and missing masks, and metric computation omits outputs for failed sensors. TABLE VI S ENSOR -FAILURE A DAPTABILITY S TUDY ON P E MS-SD. Model

12.5% (8/64) Failed

25% (16/64) Failed

50% (32/64) Failed

MAE

MSE

MRE

MAE

MSE

MRE

MAE

MSE

MRE

BRITS T IMES N ET P O G E VON

22.770 43.377 19.346

1430.077 4889.382 779.665

0.057 0.111 0.051

22.868 54.791 19.911

1443.129 7017.141 824.927

0.062 0.140 0.052

23.910 106.984 22.650

1532.279 17557.797 1021.497

0.066 0.273 0.061

A DACTS I

16.525

611.174

0.043

16.899

638.036

0.044

17.286

669.428

0.046

Table VI shows that the errors of all imputers increase with the failed-sensor ratio because less spatial information is

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

available. A DACTS I consistently achieves the lowest MAE, MSE, and MRE across all ratios. P O G E VON remains the strongest baseline, BRITS is relatively stable, and T IMES N ET degrades more sharply as failures increase. These results show that A DACTS I preserves strong imputation accuracy under S::SFA. D. Empirical Verification for S::EIA In practical applications, only a small subset of sensors may require imputation at a given time. Full-sensor processing can therefore introduce unnecessary computation. CW-SS addresses this selective-imputation setting by augmenting the incomplete sensors with spatially informative complete sensors only when the incomplete set is smaller than a threshold K. We isolate the contribution of CW-SS under a 2% evaluation missing rate. Missing-rate sensitivity is studied separately in Section V-G. We compare four A DACTS I variants: A DACTS I -O NLY uses only incomplete sensors; A DACTS I R AND adds random complete sensors up to K; A DACTS I -CW selects additional sensors by CW-SS; and A DACTS I -A LL uses all sensors as a full-sensor reference. We set K = 32 for the selective variants. The table also lists P O G E VON under its standard full-sensor setting. TABLE VII CW-SS S ENSOR -S ELECTION A BLATION . PeMS-BA

Model

PeMS-LA

PeMS-SD

MAE

MSE

MRE

MAE

MSE

MRE

MAE

MSE

MRE

P O G E VON A DACTS I -A LL

20.838 17.384

1177.320 941.174

0.052 0.043

22.706 20.626

1510.112 1389.273

0.043 0.039

17.112 14.016

742.957 564.885

0.044 0.036

A DACTS I -O NLY A DACTS I -R AND A DACTS I -CW

20.717 20.043 19.512

1181.736 1121.233 1081.589

0.052 0.050 0.049

23.293 22.229 22.163

1601.004 1471.157 1470.742

0.044 0.042 0.042

17.524 16.654 16.376

783.050 716.679 693.495

0.045 0.043 0.042

Table VII shows that A DACTS I -A LL achieves the lowest error with full-sensor input. Among the selective variants, A DACTS I -CW performs best on all three PeMS datasets, outperforming both A DACTS I -O NLY and A DACTS I -R AND. It also attains lower MAE than P O G E VON under full-sensor input, showing the value of selecting informative spatial context for selective imputation. This confirms the specific contribution of CW-SS to S::EIA. E. Empirical Verification for S::DRA In real-world applications, imputers must operate across devices with varying memory and processing capacities. To evaluate how A DACTS I addresses such computational mismatches on demand, we focus on three aspects: adaptation to memory capacity, adaptation to processing power, and the balance between input workload and model complexity. 1) Adaptation to Memory Capacity: Devices such as smart meters, wearables, and underwater monitors often have limited and varying memory capacities. To assess how A DACTS I adapts to different memory resources on demand, we simulate memory constraints by partitioning the input CTS data along the sensor dimension. Specifically, we divide the input window into g = 1, 2, 4, 8, 16 groups, each containing an equal number of sensors. These partitions are sequentially fed into the trained A DACTS I model, simulating devices with different memory

10

capacities. As illustrated in Fig. 6 (a)–(c), increasing the number of partition groups slightly degrades accuracy but keeps A DACTS I competitive with strong full-input baselines. We include the fixed full-input accuracies of P O G E VON, BRITS, and T IMES N ET from Table IV as horizontal references for this sensor-partition trade-off. As illustrated in Fig. 6 (d), increasing the number of groups reduces the per-pass sensor-partition size and lowers peak memory usage. This partitioning is a resource-control mechanism rather than an accuracy-enhancement step: smaller partitions reduce memory but slightly increase MAE because spatial correlations can be captured only within each processed sensor group. Nevertheless, performance remains competitive with baseline methods. Notably, when g = 16, memory demand falls below 2 MB, enabling A DACTS I to be deployed on MCUs such as the STM32H7 [3], [28], [46]. This result demonstrates a favorable trade-off between memory usage and accuracy. The ability to adjust memory demand enhances A DACTS I’s usability on resource-constrained devices and broadens its application range. 2) Adaptation to Processing Power: To evaluate the adaptability of A DACTS I to different levels of processing power, we adjust the attention sparsity, which determines the number of sensors used in sparse spatial attention and directly influences computational demand. By setting the attention sparsity rate to 0%, 25%, 50%, or 75%, we simulate environments with different processing capacities. A higher sparsity rate excludes a larger proportion of sensors from self-attention, thereby reducing computational demand. TABLE VIII ACCURACY AT D IFFERENT ATTENTION S PARSITY R ATES . PeMS-BA

Sparsity

PeMS-LA

PeMS-SD

MAE

MSE

MRE

MAE

MSE

MRE

MAE

MSE

MRE

P O G E VON

22.194

1248.68

0.055

23.905

1714.96

0.045

18.990

951.56

0.048

A DACTS I-75% A DACTS I-50% A DACTS I-25% A DACTS I-0%

22.741 20.745 19.505 19.433

1287.47 1176.77 1061.09 1056.99

0.056 0.050 0.048 0.048

23.109 22.801 22.565 22.143

1762.38 1784.14 1774.18 1654.58

0.044 0.043 0.042 0.041

18.150 17.950 16.142 15.945

906.05 889.47 767.30 754.23

0.046 0.046 0.041 0.041

Table VIII demonstrates that A DACTS I maintains acceptable imputation accuracy as the attention sparsity rate increases and processing demand decreases. Using all sensors (i.e., a sparsity rate of 0%) achieves the highest accuracy but requires spatial attention over all input sensors and thus greater processing power. At a sparsity rate of 75%, SSA uses only 25% of the sensors and incurs only a modest performance decrease, remaining comparable to the strongest baseline, P O G E VON (e.g., MAE values of 22.741 vs. 22.194 on PeMSBA). These findings show that A DACTS I can balance processing demand and imputation accuracy on demand. This adaptability makes A DACTS I well suited to elastic computing applications, in which processing power scales with resource availability. 3) Balancing Input Workload with Model Complexity: As the number of input sensors—and hence the workload— varies, adjusting model complexity is essential for optimizing performance under stable but limited resource conditions. In this experiment, we explore how A DACTS I balances input

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

POGEVON

ADACTSI

Peak Mem (MB)

MAE

Peak Mem Usage of AdaCTSi

22

32

26

TIMESNET

BRITS

36

30

11

20

28

18

22 24

16

Partition Group Numbers

2

0

1

2

2

2

3

2

2

4

(a) MAE on PeMS-BA

2

0

2

1

2

2

2

3

2

4

20

(b) MAE on PeMS-LA

21

22

23

8

Inadaptive

6 4 2

20

24

21

22

23

24

(d) Memory Usage

(c) MAE on PeMS-SD

Fig. 6. MAE and peak memory usage of A DACTS I with different numbers of sensor-partition groups (g = 1, 2, 4, 8, 16). Horizontal baseline lines are fixed full-input references from Table IV.

8 75% 22.15

Sensor Partition Size 16 32 20.59 18.82

64 18.15

19.46

18.05

17.95

25% 21.46

19.03

17.50

16.68

0% 21.25

18.39

16.53

15.95

(MB)

50% 21.70

Mem

Sparsity Rate

workload and model complexity on demand to maintain high accuracy while satisfying fixed resource constraints. We vary the input workload by adjusting the sensor-partition size and modify model complexity by altering attention sparsity, following the settings used in the preceding experiments. This approach allows us to observe how different combinations of workload and complexity affect model accuracy while resource demands remain stable.

8.68 2.24

3.16

5.00

baseline, in which missing values are filled with zeros, to evaluate the benefits of learned imputation. Each method fills missing values before forecasting, ensuring a fair and consistent comparison of its effect on forecasting accuracy. Using the PeMS-SD dataset, we partition the data into training and testing sets at a 60:40 ratio. We explore two typical historical window sizes: a short-term window of 3 time steps (15 minutes) and a long-term window of 24 time steps (2 hours). In both cases, we forecast the next time step (5 minutes ahead). Forecasting accuracy is evaluated against the ground-truth values following established practices in CTS forecasting [27], [28]. For forecasting, we employ widely used models, including Random Forest (RF) [47], Gradient Boosting (GB) [48], XGBoost [49], LightGBM [50], and Decision Tree (DT) [51]. These models provide diverse perspectives on the quality of the imputed data. TABLE IX D OWNSTREAM F ORECASTING R ESULTS ON P E MS-SD.

Fig. 7. Trade-off Between Input Workload (i.e., Sensor Partition Size) and Model Complexity (i.e., Sparsity Rate).

Forecasting

As illustrated in Fig. 7, A DACTS I balances input workload and model complexity to maintain robust on-demand performance. Each block represents a combination of input partition size (workload) and attention sparsity rate (model complexity). The primary factor determining peak memory usage is the input workload because a larger number of input sensors increases memory demand. Consequently, A DACTS I adjusts its computational demands in response to online workloads, ensuring efficient resource utilization while sustaining high imputation accuracy. Overall, these experiments demonstrate that A DACTS I can adapt on demand to online environments with varying memory capacities, levels of processing power, and combinations of input workload and model complexity. This adaptability enables A DACTS I to operate efficiently across a wide range of devices and settings (S::DRA) while maintaining high imputation accuracy and effective resource utilization. F. Verification Using Downstream Forecasting To evaluate the effectiveness of A DACTS I in improving data quality, we conduct one-step-ahead forecasting experiments. Specifically, we compare A DACTS I against the top three baselines, P O G E VON, T IMES N ET, and BRITS, as identified in Section V-B. We also include a No-Imputation

Imputation

Historical Window Size = 3

Historical Window Size = 24

MAE

MSE

MAPE

MAE

MSE

MAPE

RF

No Imputation BRITS T IMES N ET P O G E VON A DACTS I

23.519 21.854 21.207 20.874 20.642

1100.674 925.012 838.800 808.613 793.155

9.36% 8.25% 8.69% 8.34% 8.29%

26.072 24.171 21.854 21.011 20.597

1358.583 1115.385 880.299 819.290 786.917

9.87% 8.88% 8.89% 8.36% 8.22%

GB

No Imputation BRITS T IMES N ET P O G E VON A DACTS I

28.711 21.424 21.307 21.118 20.696

1560.026 858.889 842.096 822.278 791.626

11.69% 8.12% 8.92% 8.59% 8.63%

30.593 22.271 21.247 20.956 20.584

1824.499 924.938 839.423 817.093 795.069

11.96% 8.54% 8.88% 8.43% 8.38%

XGBoost

No Imputation BRITS T IMES N ET P O G E VON A DACTS I

27.644 24.631 22.727 22.061 22.005

1482.383 1144.539 998.016 928.763 931.100

11.46% 9.28% 9.13% 8.72% 8.69%

29.800 25.047 23.153 22.783 22.634

1791.796 1141.652 1077.440 1015.970 1000.774

11.33% 9.41% 9.09% 8.78% 8.79%

LightGBM

No Imputation BRITS T IMES N ET P O G E VON A DACTS I

24.650 22.080 20.938 20.324 20.146

1180.375 897.033 826.105 762.179 752.057

10.26% 8.86% 8.59% 8.15% 8.11%

26.091 23.492 20.926 20.275 19.637

1358.139 1019.098 819.107 771.485 709.857

9.90% 8.91% 8.57% 8.08% 7.96%

DT

No Imputation BRITS T IMES N ET P O G E VON A DACTS I

37.188 34.201 30.776 30.882 29.408

3298.900 2172.071 1829.297 1756.827 1643.930

14.79% 12.68% 12.67% 12.53% 11.65%

41.742 34.083 31.921 31.487 29.816

3803.346 2439.364 1919.124 1918.582 1730.376

15.50% 13.20% 13.28% 12.38% 11.83%

Table IX demonstrates that A DACTS I generally outperforms the baseline methods across forecasting models and historical window sizes. In most cases, A DACTS I achieves the lowest MAE, MSE, and MAPE, with particularly notable improvements for LightGBM and Random Forest. In these cases, it surpasses leading approaches such as P O G E VON, T IMES N ET, and BRITS, demonstrating its effectiveness across diverse forecasting scenarios.

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

PeMS-BA

PeMS-LA

PeMS-SD

MAE

60 40 20 15

45

65

85 15

ADACTSI

45

65

85 15

NET3

BRITS

Missing rate (%)

POGEVON

45

65

85

RGAIN

Fig. 8. Impact of missing rate.

A longer historical window provides more temporal information but does not consistently improve forecasting accuracy, potentially because of overfitting or the inclusion of outdated information. This observation underscores the importance of high-quality CTS imputation rather than merely increasing data volume; effective imputers such as A DACTS I support accurate forecasts across a range of models and configurations. These results confirm that the accurate imputations produced by A DACTS I yield strong downstream forecasting performance. By providing reliable inputs, A DACTS I helps forecasting models generate better predictions, underscoring its practical value in applications with incomplete data. G. Robustness Study To evaluate the robustness of A DACTS I under different levels of data scarcity, we conduct a sensitivity analysis on PeMS-BA, PeMS-LA, and PeMS-SD using missing rates from 15% to 85% in 10-percentage-point increments. As illustrated in Fig. 8, the MAE of all methods generally increases as the missing rate grows, reflecting the inherent difficulty of imputation when fewer observations are available. A DACTS I consistently achieves the lowest MAE, reducing error by 7.1% on average compared with the strongest baseline. Even at high missing rates of 55%–85%, it reduces MAE by 5.4% on average. This trend demonstrates the effectiveness and robustness of our adaptive architecture under increasingly sparse observations. H. Efficiency Study We evaluate efficiency through offline training and online inference costs. Table X reports PeMS results under a common training protocol.

12

profiling [52]; the dash for NET3 denotes a torchinfo incompatibility. A DACTS I combines adaptive sensor-subset and resource-aware inference with competitive efficiency. Its training cost is close to lightweight deep baselines and much lower than graph-heavy or generative models. For online deployment, A DACTS I achieves the lowest inference latency among the deep models while keeping a modest memory footprint. These results support the intended deployment pattern: offline training can be completed once on a server, and the resulting model can be reused across resource settings. I. Ablation Studies To evaluate the contribution of each component in A DACTS I, we conduct ablation studies on the three PeMS datasets. By systematically removing key components, we assess the effect of each module on imputation performance. Specifically, we create the following variants: A DACTS I LTSIT removes LTSIT, leaving the One-shot TCN to process the raw input without learned multi-view embeddings; A DACTS I -OSTCN replaces the One-shot TCN with a standard MLP that lacks temporal convolutions; and A DACTS I SSA removes the SSA module for spatial correlation extraction. The quantitative results in Table XI show that the complete A DACTS I architecture outperforms all ablated variants, demonstrating the contribution of each module. The higher errors of A DACTS I -LTSIT confirm the role of LTSIT in learning multi-view embeddings. Similarly, the degradation of A DACTS I -OSTCN demonstrates the importance of the temporal convolutions in the One-shot TCN, while the gap between A DACTS I -SSA and the full model confirms the value of SSA for capturing spatial correlations among sensors. Overall, the ablation studies validate the contribution of each component in A DACTS I. The full model outperforms all ablated variants, showing that integrating LTSIT, One-shot TCN, and SSA improves imputation performance. TABLE XI A BLATION S TUDY R ESULTS . PeMS-BA

Model

PeMS-LA

PeMS-SD

MAE

MSE

MRE

MAE

MSE

MRE

MAE

MSE

MRE

A DACTS I -LTSIT A DACTS I -OSTCN A DACTS I -SSA

21.271 26.412 20.234

1124.562 1812.084 1081.424

0.053 0.065 0.050

26.509 27.675 23.890

1899.853 2104.829 1685.953

0.055 0.057 0.050

19.925 24.518 16.599

1039.318 1716.217 803.075

0.051 0.063 0.042

A DACTS I

19.433

1056.990

0.048

22.141

1654.584

0.041

15.945

754.236

0.041

TABLE X E FFICIENCY C OMPARISON ON P E MS. Method

Training cost

Inference cost

J. Parameter Studies

Time (h) Peak Mem. (MB) Lat. (ms) Peak Mem. (MB) P O G E VON NET3 GRIN SAITS BRITS R GAIN T IMES N ET

39.62 65.49 23.01 4.00 7.52 3.47 4.43

10867.5 593.8 919.4 83.0 61.6 84.2 75.3

172.82 130.19 127.81 45.23 40.65 18.82 10.71

A DACTS I

3.71

249.1

9.57

16.08 12.62 7.35 2.69 5.02 6.24 8.68

Table X reports training memory as peak CUDA allocation and inference memory using batch-1 torchinfo

We investigate the influence of three key parameters—the embedding size D, window size T, and CW threshold K— on imputation performance. We use MAE on the PeMS-SD dataset under the experimental configurations described in Sections V-A and V-D. Embedding Size D: As illustrated in Fig. 9 (left), increasing the embedding size D from 72 to 144 consistently reduces the MAE, reflecting improved imputation accuracy. The optimal embedding size is observed at D = 144, which effectively balances model expressiveness and regularization, thereby

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

minimizing the risk of overfitting. However, further increasing D beyond 144 results in a slight degradation in performance. This decline indicates that excessively large embeddings may introduce redundant parameters, leading to increased computational complexity, potential overfitting, and diminished generalization capability. Window Size T: The impact of varying T is presented in Fig. 9 (middle). As the window size increases from 12 to 24, MAE decreases, indicating a benefit from the richer temporal context. The best performance is achieved at T = 24, which represents two hours of data at five-minute sampling intervals. Window sizes of 12, 18, and 24 divide evenly into the 288step daily cycle, whereas a window size of 30 does not and yields a higher MAE. These results suggest that both temporal context and alignment with the data’s periodic structure affect model performance. CW Threshold K: The sensitivity of the model to K is depicted in Fig. 9 (right). MAE stabilizes when K ≥ 32. Given a total of 64 sensors, setting K = 32 ensures the selection of at least half of the sensors, providing adequate spatial information for accurate imputation. Thresholds below 32 select fewer sensors, limit the available spatial context, and increase MAE. We therefore recommend setting the CW threshold to at least half of the sensor count.

17

MAE

MAE

21

17

19

16

MAE vs CW Threshold

18

23

18

15

MAE vs Window Size

MAE vs Embedding Size

MAE

19

16

17 72

96

120

144

Embedding Size (D)

168

15

12

18

24

Window Size (T)

30

15

8

16

24

32

CW Threshold (K)

40

Fig. 9. Impact of D, T, and K.

Overall, A DACTS I exposes only a small set of key parameters through its streamlined architecture. One-shot TCN, LTSIT, SSA, and CW-SS are each used once rather than repeated across stacked modules. This design simplifies model configuration and deployment across diverse environments. VI. R ELATED W ORK Deep Models for CTS Imputation. CTS imputation has been studied through statistical, matrix-factorization, subspace, and DL-based approaches. Representative classical baselines include MEAN, Matrix Factorization (MF), and Multivariate Imputation by Chained Equations (MICE) [35]. Classical matrix-completion and subspace-based methods introduce structured priors through spectral regularization for lowrank matrix completion [53], streaming pattern discovery in multivariate time series [54], temporally regularized matrix factorization (TRMF) [36], ST-MVL for geo-sensory time series [55], spatial multi-view NMF (SMV-NMF) for sensor and urban statistical data [37], and scalable recovery of missing blocks under different cross-correlation regimes [56]. These approaches remain useful when their statistical or lowrank assumptions match the data, whereas modern largescale CTS deployments often involve dynamic spatio-temporal

13

correlations and heterogeneous missing patterns that benefit from representation learning. DL models capture such complex patterns using modules such as the bidirectional RNNs in BRITS [32], which incorporate temporal information from both past and future timestamps. More recent models enhance spatial feature extraction using GNNs [20], [40], [41] and self-attention mechanisms [19], thereby capturing both spatial and temporal relationships [39]. Generative models, including GANs [38], [57] and VAEs [16], [58], learn latent representations to generate imputed values. Diffusion models [14], [15], [59] further improve generative modeling through iterative sampling and are often studied in offline imputation settings. R E CTS I [18] studies efficient CTS imputation under a fixed input and model configuration. Most of these imputers are evaluated with fixed sensor sets and model configurations. In contrast, A DACTS I uses one trained model for elastic inference across variable sensor subsets and resource budgets. It combines One-shot TCN and LTSIT for sensor-wise representations, SSA for spatial correlation extraction, and CW-SS for informative-sensor selection. Adaptability in DL Methods. Adaptability is essential for deploying DL models in environments with varying computational resources and input workloads. It requires balancing computational efficiency with accuracy, handling diverse workloads, and supporting devices with different resource capacities. Existing solutions predominantly maintain a collection of models and select one based on the input workload and available resources. Although this approach is adaptable, designing, training, and updating multiple models incurs substantial computational and storage overhead. In computer vision, some methods use patch-wise processing [60]–[62], dividing images into patches to handle varying image sizes by leveraging local spatial characteristics and texture patterns. This idea is less directly applicable to CTS imputation because CTS workloads comprise disjoint time series subsets with distinct characteristics and patterns [28]. In summary, although model collections and patch-wise processing can enhance adaptability, they introduce modelmanagement complexity and do not directly accommodate disjoint CTS subsets. To address these issues, we propose A DACTS I, an adaptive CTS imputer for changing environments. A DACTS I features an On-demand Mode that uses a single adaptable N -in-1 model across varying sensor subsets and resource settings. VII. C ONCLUSION AND F UTURE W ORK In this paper, we presented A DACTS I, an adaptive and efficient CTS imputer for changing environments. It addresses sensor failures, excessive imputation, and resource variability by dynamically adjusting its input workload and computational requirements, thereby establishing a foundation for adaptive CTS imputation. In future work, we will incorporate predictive resource provisioning and scheduling to anticipate workload and resource changes, enhance on-demand adaptability, and validate these

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

improvements in real-world online scenarios. This direction involves combining workload and resource prediction with dynamic cluster adjustment, which is beyond the current scope but important for complex, real-time IoT applications. We also aim to integrate lightweight modules to improve A DACTS I’s efficiency and enable faster on-demand CTS imputation. ACKNOWLEDGMENT This work was supported by the Fujian Provincial Artificial Intelligence Industry Development Technology Project under Grant 2025H0042 and by the Young Scientists Fund of the National Natural Science Foundation of China under Grant 62402420 for the project “Research on Key Technologies for Edge-Intelligent Time-Series Data Imputation.” R EFERENCES [1] K. Zhao, C. Guo, Y. Cheng, P. Han, M. Zhang, and B. Yang, “Multiple time series forecasting with dynamic graph modeling,” Proc. VLDB Endow., vol. 17, no. 4, pp. 753–765, 2023. [2] Z. Lai, D. Zhang, H. Li, C. S. Jensen, H. Lu, and Y. Zhao, “LightCTS*: lightweight correlated time series forecasting enhanced with model distillation,” IEEE Trans. Knowl. Data Eng., vol. 36, no. 12, pp. 8695– 8710, 2024. [3] Z. Lai, H. Li, D. Zhang, Y. Zhao, W. Qian, and C. S. Jensen, “E2Usd: efficient-yet-effective unsupervised state detection for multivariate time series,” in Proc. ACM Web Conf., 2024, pp. 3010–3021. [4] C. Wang, K. Wu, T. Zhou, and Z. Cai, “Time2State: an unsupervised framework for inferring the latent states in time series data,” Proc. ACM Manag. Data, vol. 1, no. 1, pp. 1–18, 2023. [5] Y. Yao, D. Li, H. Jie, T. Li, J. Chen, J. Wang, F. Li, and Y. Gao, “SimpleTS: an efficient and universal model selection framework for time series forecasting,” Proc. VLDB Endow., vol. 16, no. 12, pp. 3741– 3753, 2023. [6] H. Mohr-Daurat, G. Theodorakis, and H. Pirk, “Hardware-efficient data imputation through DBMS extensibility,” Proc. VLDB Endow., vol. 17, no. 11, pp. 3497–3510, 2024. [7] A. Khelifati, M. Khayati, A. Dignös, D. Difallah, and P. CudréMauroux, “TSM-Bench: benchmarking time series database systems for monitoring applications,” Proc. VLDB Endow., vol. 16, no. 11, pp. 3363– 3376, 2023. [8] C. Wang, X. Huang, J. Qiao, T. Jiang, L. Rui, J. Zhang, R. Kang, J. Feinauer, K. A. McGrail, P. Wang et al., “Apache IoTDB: time-series database for the Internet of Things,” Proc. VLDB Endow., vol. 13, no. 12, pp. 2901–2904, 2020. [9] D. Campos, M. Zhang, B. Yang, T. Kieu, C. Guo, and C. S. Jensen, “LightTS: lightweight time series classification with adaptive ensemble distillation,” Proc. ACM Manag. Data, vol. 1, no. 2, pp. 1–27, 2023. [10] X. He, Y. Li, J. Tan, B. Wu, and F. Li, “OneShotSTL: one-shot seasonal-trend decomposition for online time series anomaly detection and forecasting,” Proc. VLDB Endow., vol. 16, no. 6, pp. 1399–1412, 2023. [11] D. Campos, T. Kieu, C. Guo, F. Huang, K. Zheng, B. Yang, and C. S. Jensen, “Unsupervised time series outlier detection with diversity-driven convolutional ensembles,” Proc. VLDB Endow., vol. 15, no. 3, pp. 611– 623, 2021. [12] L. Chen, D. Chen, Z. Shang, B. Wu, C. Zheng, B. Wen, and W. Zhang, “Multi-scale adaptive graph neural network for multivariate time series forecasting,” IEEE Trans. Knowl. Data Eng., vol. 35, no. 10, pp. 10 748– 10 761, 2023. [13] A. Blázquez-Garcı́a, K. Wickstrøm, S. Yu, K. Ø. Mikalsen, A. Boubekki, A. Conde, U. Mori, R. Jenssen, and J. A. Lozano, “Selective imputation for multivariate time series datasets with missing values,” IEEE Trans. Knowl. Data Eng., vol. 35, no. 9, pp. 9490–9501, 2023. [14] Y. Tashiro, J. Song, Y. Song, and S. Ermon, “CSDI: conditional scorebased diffusion models for probabilistic time series imputation,” in Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 24 804–24 816. [15] M. Liu, H. Huang, H. Feng, L. Sun, B. Du, and Y. Fu, “PriSTI: a conditional diffusion framework for spatiotemporal imputation,” in Proc. IEEE 39th Int. Conf. Data Eng. (ICDE), 2023, pp. 1927–1939.

14

[16] D. Wang, Y. Yan, R. Qiu, Y. Zhu, K. Guan, A. Margenot, and H. Tong, “Networked time series imputation via position-aware graph enhanced variational autoencoders,” in Proc. ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2023, pp. 2256–2268. [17] H. Yuan, G. Cong, and G. Li, “Nuhuo: an effective estimation model for traffic speed histogram imputation on a road network,” Proc. VLDB Endow., vol. 17, no. 7, pp. 1605–1617, May 2024. [18] Z. Lai, D. Zhang, H. Li, D. Zhang, H. Lu, and C. S. Jensen, “ReCTSi: resource-efficient correlated time series imputation via decoupled pattern learning and integrity-aware attentions,” in Proc. ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2024, pp. 1474–1483. [19] W. Du, D. Côté, and Y. Liu, “SAITS: self-attention-based imputation for time series,” Expert Syst. Appl., vol. 219, p. 119619, 2023. [20] X. Li, H. Li, H. Lu, C. S. Jensen, V. Pandey, and V. Markl, “Missing value imputation for multi-attribute sensor data streams via message propagation,” Proc. VLDB Endow., vol. 17, no. 3, pp. 345–358, 2023. [21] Y. Li, Y. Lin, Y. Wang, K. Ye, and C. Xu, “Serverless computing: stateof-the-art, challenges and opportunities,” IEEE Trans. Services Comput., vol. 16, no. 2, pp. 1522–1539, 2022. [22] S. Fouladi, R. S. Wahby, B. Shacklett, K. V. Balasubramaniam, W. Zeng, R. Bhalerao, A. Sivaraman, G. Porter, and K. Winstein, “Encoding, fast and slow: low-latency video processing using thousands of tiny threads,” in Proc. USENIX Symp. Netw. Syst. Design Implementation (NSDI), 2017, pp. 363–376. [23] Z. Li, L. Guo, J. Cheng, Q. Chen, B. He, and M. Guo, “The serverless computing survey: a technical primer for design architecture,” ACM Comput. Surv., vol. 54, no. 10s, pp. 1–34, 2022. [24] H. Shen and L. Chen, “A resource-efficient predictive resource provisioning system in cloud systems,” IEEE Trans. Parallel Distrib. Syst., vol. 33, no. 12, pp. 3886–3900, 2022. [25] O. Poppe, T. Amuneke, D. Banda, A. De, A. Green, M. Knoertzer, E. Nosakhare, K. Rajendran, D. Shankargouda, M. Wang et al., “Seagull: an infrastructure for load prediction and optimized resource allocation,” Proc. VLDB Endow., vol. 14, no. 2, pp. 154–162, 2020. [26] N. Bashir, N. Deng, K. Rzadca, D. Irwin, S. Kodak, and R. Jnagal, “Take it to the limit: peak prediction-driven resource overcommitment in datacenters,” in Proc. ACM Eur. Conf. Comput. Syst. (EuroSys), 2021, pp. 556–573. [27] X. Wu, D. Zhang, C. Guo, C. He, B. Yang, and C. S. Jensen, “AutoCTS: automated correlated time series forecasting,” Proc. VLDB Endow., vol. 15, no. 4, pp. 971–983, 2021. [28] Z. Lai, D. Zhang, H. Li, C. S. Jensen, H. Lu, and Y. Zhao, “LightCTS: a lightweight framework for correlated time series forecasting,” Proc. ACM Manag. Data, vol. 1, no. 2, pp. 1–26, 2023. [29] F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2016, pp. 1–10. [30] R. Qian, S. Ding, X. Liu, and D. Lin, “Static and dynamic concepts for self-supervised video representation learning,” in Proc. Eur. Conf. Comput. Vis. (ECCV), 2022, pp. 145–164. [31] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inf. Process. Syst., vol. 30, 2017, pp. 5998–6008. [32] W. Cao, D. Wang, J. Li, H. Zhou, L. Li, and Y. Li, “BRITS: bidirectional recurrent imputation for time series,” in Adv. Neural Inf. Process. Syst., vol. 31, 2018, pp. 6776–6786. [33] R. J. Little and D. B. Rubin, Statistical Analysis With Missing Data, 3rd ed. Hoboken, NJ, USA: Wiley, 2019. [34] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed. New York, NY, USA: Springer, 2009. [35] I. R. White, P. Royston, and A. M. Wood, “Multiple imputation using chained equations: issues and guidance for practice,” Stat. Med., vol. 30, no. 4, pp. 377–399, 2011. [36] H.-F. Yu, N. Rao, and I. S. Dhillon, “Temporal regularized matrix factorization for high-dimensional time series prediction,” in Adv. Neural Inf. Process. Syst., vol. 29, 2016, pp. 847–855. [37] Y. Gong, Z. Li, J. Zhang, W. Liu, B. Chen, and X. Dong, “A spatial missing value imputation method for multi-view urban statistical data,” in Proc. 29th Int. Joint Conf. Artif. Intell. (IJCAI), 2020, pp. 1310–1315. [38] X. Miao, Y. Wu, J. Wang, Y. Gao, X. Mao, and J. Yin, “Generative semisupervised learning for multivariate time series imputation,” in Proc. AAAI Conf. Artif. Intell., vol. 35, no. 10, 2021, pp. 8983–8991. [39] H. Wu, T. Hu, Y. Liu, H. Zhou, J. Wang, and M. Long, “TimesNet: temporal 2D-variation modeling for general time series analysis,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2023.

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, VOL. XX, NO. XX, XX 20XX

[40] A. Cini, I. Marisca, and C. Alippi, “Filling the gaps: multivariate time series imputation by graph neural networks,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2022. [Online]. Available: https: //openreview.net/forum?id=kOu3-S3wJ7 [41] B. Jing, H. Tong, and Y. Zhu, “Network of tensor time series,” in Proc. ACM Web Conf., 2021, pp. 2425–2437. [42] “AdaCTSi project: code, datasets, and supplemental material,” 2025, accessed: Jul. 2026. [Online]. Available: https://github.com/ryanlaics/ adactsi/ [43] C. Chen, K. Petty, A. Skabardonis, P. Varaiya, and Z. Jia, “Freeway performance measurement system: mining loop detector data,” Transp. Res. Rec., vol. 1748, no. 1, pp. 96–102, 2001. [44] S. Zhang, B. Guo, A. Dong, J. He, Z. Xu, and S. X. Chen, “Cautionary tales on air-quality improvement in Beijing,” Proc. R. Soc. A, vol. 473, no. 2205, p. 20170457, 2017. [45] L. Grgičević, “AIS data,” Zenodo, version 1.0.1, Jun. 2023. [Online]. Available: https://doi.org/10.5281/zenodo.8064564 [46] STMicroelectronics, “Microcontrollers and microprocessors,” 2023, accessed: Jul. 2026. [Online]. Available: https://www.st.com/en/ microcontrollers-microprocessors.html [47] L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001. [48] J. H. Friedman, “Greedy function approximation: a gradient boosting machine,” Ann. Statist., vol. 29, no. 5, pp. 1189–1232, 2001. [49] T. Chen and C. Guestrin, “XGBoost: a scalable tree boosting system,” in Proc. ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2016, pp. 785–794. [50] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, “LightGBM: a highly efficient gradient boosting decision tree,” in Adv. Neural Inf. Process. Syst., vol. 30, 2017, pp. 3146–3154. [51] L. Breiman, Classification and Regression Trees. New York, NY, USA: Routledge, 2017. [52] T. Yep, “Torchinfo,” 2020, accessed: Jul. 2026. [Online]. Available: https://github.com/TylerYep/torchinfo [53] R. Mazumder, T. Hastie, and R. Tibshirani, “Spectral regularization algorithms for learning large incomplete matrices,” J. Mach. Learn. Res., vol. 11, no. 80, pp. 2287–2322, 2010. [54] S. Papadimitriou, J. Sun, and C. Faloutsos, “Streaming pattern discovery in multiple time-series,” in Proc. 31st Int. Conf. Very Large Data Bases (VLDB), 2005, pp. 697–708. [55] X. Yi, Y. Zheng, J. Zhang, and T. Li, “ST-MVL: filling missing values in geo-sensory time series data,” in Proc. 25th Int. Joint Conf. Artif. Intell. (IJCAI), 2016, pp. 2710–2716. [56] M. Khayati, P. Cudré-Mauroux, and M. H. Böhlen, “Scalable recovery of missing blocks in time series with high and low cross-correlations,” Knowl. Inf. Syst., vol. 62, no. 6, pp. 2275–2308, 2020. [57] Y. Luo, Y. Zhang, X. Cai, and X. Yuan, “E2GAN: end-to-end generative adversarial network for multivariate time series imputation,” in Proc. Int. Joint Conf. Artif. Intell. (IJCAI), 2019, pp. 3094–3100. [58] V. Fortuin, D. Baranchuk, G. Rätsch, and S. Mandt, “GP-VAE: deep probabilistic time series imputation,” in Proc. Int. Conf. Artif. Intell. Statist. (AISTATS), 2020, pp. 1651–1661. [59] X. Wang, H. Zhang, P. Wang, Y. Zhang, B. Wang, Z. Zhou, and Y. Wang, “An observed value consistent diffusion model for imputing missing values in multivariate time series,” in Proc. ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2023, pp. 2409–2418. [60] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021. [61] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 10 012–10 022. [62] Q. Zhang and Y.-B. Yang, “ResT: an efficient transformer for visual recognition,” in Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 15 475–15 485.

15

Zhichen Lai received the bachelor’s and master’s degrees in Computer Science from the University of Electronic Science and Technology of China and Sichuan University, China, in 2018 and 2021, respectively, and the PhD degree in Computer Science from Aalborg University, Denmark, in 2025. He is currently an Associate Research Professor with the College of Computer and Data Science, Fuzhou University, China. In 2025, he was selected for China’s Overseas Talent Recruitment Program. His research focuses on spatio-temporal data mining. Huan Li (Member, IEEE) received the PhD degree from Zhejiang University in 2018. He was an assistant professor at Aalborg University in Denmark from 2020 to 2023 and a senior engineer at Alibaba from 2018 to 2019. He is currently a ZJU100 professor at Zhejiang University and was awarded an EU Marie Curie Individual Fellowship. His work has appeared in top-tier journals and conferences. His research interests include data management for AI, spatio-temporal computing, and mobile computing.

Dalin Zhang (Senior Member, IEEE) received the BS degree from Jilin University in 2012, the MS degree from the University of Chinese Academy of Sciences in 2015, and the PhD degree from the University of New South Wales (UNSW Sydney) in 2020. He is currently a professor at Hangzhou Dianzi University, China. His research interests include brain–computer interfaces, human activity recognition, and the Internet of Things.

Dong Gong is a Senior Lecturer and ARC DECRA Fellow at the School of Computer Science and Engineering, The University of New South Wales (UNSW Sydney), Australia. He also holds an adjunct position with the Australian Institute for Machine Learning (AIML) at The University of Adelaide. Before joining UNSW in 2022, he was a Research Fellow at AIML and a Principal Researcher at the Centre for Augmented Reasoning (CAR), The University of Adelaide.

Lina Yao (Senior Member, IEEE) received the master’s and Ph.D. degrees in computer science from The University of Adelaide, Australia, in 2010 and 2014, respectively. She is a Professor at UNSW Sydney and an ARC Future Fellow, as well as a Visiting Scientist at CSIRO’s Data61 and a Co-Theme Leader at the Responsible AI Research Centre. She has been a Clarivate Highly Cited Researcher since 2024. Her research interests include machine learning and data mining for the Internet of Things, recommendation, human activity recognition, brain–computer interfaces, and causal and responsible AI. Christian S. Jensen (Fellow, IEEE) received the PhD degree from Aalborg University in 1991 and the DrTechn degree from Aalborg University in 2000. He is currently a professor with the Department of Computer Science, Aalborg University. His research focuses primarily on temporal and spatio-temporal data management and analytics, including indexing and query processing, data mining, and machine learning.

Record · ID 411161 · SHA-256 56e5717ba5b5fa4e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.