ConceptioArchivearXiv CS
arXiv CSopen access

A temporal deep learning framework for calibration of low-cost air quality sensors

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

A temporal deep learning framework for calibration of low-cost air quality sensors Arindam Sengupta1,* , Tony Bush2 , Ben Marner2 , Jose Miguel Pérez1 , and Soledad Le Clainche1

arXiv:2604.21527v1 [cs.LG] 23 Apr 2026

1

ETSI Aeronáutica y del Espacio, Universidad Politécnica de Madrid, Plaza Cardenal Cisneros, 3, Madrid, 28040, Spain 2 Air Quality Consultants Ltd., 3rd Floor, St. Augustine’s Court, St. Augustine’s Place,

Bristol, BS1 4UD, United Kingdom * Corresponding author: [email protected] (Arindam Sengupta) , Co-authors:

[email protected], [email protected], [email protected], [email protected]

Abstract Low-cost air quality sensors (LCS) provide a practical alternative to expensive regulatory-grade instruments, making dense urban monitoring networks possible. Yet their adoption is limited by calibration challenges, including sensor drift, environmental cross-sensitivity, and variability in performance from device to device. This work presents a deep learning framework for calibrating LCS measurements of PM2.5 , PM10 , and NO2 using a Long Short-Term Memory (LSTM) network, trained on co-located reference data from the OxAria network in Oxford, UK. Unlike the Random Forest (RF) baseline, which treats each observation independently, the proposed approach captures temporal dependencies and delayed environmental effects through sequence-based learning, achieving higher R2 values across training, validation, and test sets for all three pollutants. A feature set is constructed combining time-lagged parameters, harmonic encodings, and interaction terms to improve generalization on unseen temporal windows. Validation of unseen calibrated values against the Equivalence Spreadsheet Tool 3.1 demonstrates regulatory compliance with expanded uncertainties of 22.11% for NO2 , 12.42% for PM10 , and 9.1% for PM2.5 .

Keywords: Low-cost sensors, air pollution, calibration, deep learning, LSTM, feature engineering.

1

1.

Introduction

Outdoor air pollution was responsible for an estimated 4.2 million premature deaths in 2019, with 89% occurring in low- and middle-income countries [1]. Fine particulate matter (PM2.5 ), coarse particulate matter (PM10 ), and nitrogen dioxide (NO2 ) are among the most harmful urban pollutants, linked to respiratory illness, cardiovascular disease, and premature mortality [1, 2]. Levels of these pollutants change sharply by location, traffic, and time of year, illustrating the dynamic nature of urban air quality. Continuous, high spatio-temporal resolution monitoring is therefore essential to accurately estimate their behaviour and guide public health measures and policies [3, 4]. However, the high cost and infrastructure requirements of reference-grade air quality monitoring stations have led to only sparse deployment [5]. In this context, low-cost sensors (LCS) have emerged as potentially viable alternatives, enabling real-time networks for urban air quality surveillance. Recent work has shown that real-time insights into urban air quality can be provided when LCSs are deployed alongside more sparse reference stations in a complementary fashion [6]. Despite their appeal, LCS measurements are often compromised by environmental cross-sensitivity, temporal drift, and inter-sensor variability [7, 8]. These limitations introduce non-linearity and systematic bias, particularly under fluctuating ambient conditions such as temperature and humidity [9]. Thus, reliable calibration is essential before these sensors can be integrated into decision-making processes or scientific studies. Calibration of low-cost air quality sensors has been approached using a range of statistical models. Regression methods remain the most widely applied, with univariate linear regression (ULR) serving as a benchmark and multivariate linear regression (MLR) improving accuracy by incorporating meteorological or co-pollutant predictors [5]. Traditional calibration of LCS has typically relied on co-location with reference-grade instruments and employed statistical methods. While effective in certain conditions, these approaches often fail to generalize across locations or over time due to their inability to capture the non-linearities and temporal dynamics inherent in sensor behaviour [5, 10]. This emphasizes the need for improved and reliable techniques for LCS calibration. Several studies have highlighted the importance of applying machine learning (ML) techniques to calibrate LCS outputs against reference instruments, reporting improved performance over traditional methods [5, 6, 10, 11]. Liang [5] provided a detailed description of various regression and ML approaches, reviewing state-of-the-art LCS calibration approaches, highlighting the shift from traditional statistical corrections to advanced data-driven models. Among these, supervised regression models such as Random Forest (RF), XGBoost, Support Vector Machines (SVM), and neural networks have been widely explored. For instance, Si et al. [10] compared MLR, XGBoost, and feed-forward neural networks for PM2.5 calibration, demonstrating that deep models consistently outperformed linear baselines, particularly under varying humidity conditions. Nowack et al.

2

[12] demonstrated the successful calibration of co-located PM10 and NO2 sensors, using methods such as ridge regression, RF regression, and Gaussian process regression (GPR). Similarly, Yin et al. [6] applied spatially-aware XGBoost models to improve generalization across urban sensor deployments. While Zimmerman et al. [13] highlighted that machine learning methods, such as RF, can effectively mitigate drift and cross-sensitivity issues. Extending this line of work, Apostolopoulos et al. [14] conducted an extensive field validation of low-cost NO2 and O3 sensors across urban deployment in Greece. They assessed machine learning approaches such as RF and LSTM models and found that incorporating meteorological context and carefully selecting input features significantly enhanced the calibration model’s ability to generalize across different environmental conditions. Spinelle et al. [15] demonstrated that supervised learning techniques performed better than traditional regression techniques. Han et al. [16] described a multi-task methodology for air quality prediction, where temporal dynamics and data preprocessing are jointly leveraged to improve forecasting performance. Mahajan and Helbing [17] introduced an innovative trust-based calibration method that dynamically adapts model complexity based on sensor reliability. They employed trust-weighted consensus and wavelet-based features to calibrate PM2.5 in real time, reducing mean absolute error by up to 68% for low-quality sensors while maintaining computational efficiency. Bigi et al. [18] and Bush et al. [19] have demonstrated that ensemble machine learning techniques, such as RF, are effective for calibrating low-cost air quality sensors. Bigi et al. [18] evaluated three calibration strategies applied across a set of electrochemical sensors located in Switzerland. Their findings demonstrated significant reductions in mean absolute error when using ensemble methods, underscoring their portability and effectiveness across diverse environments. In particular, Bush et al. [19] proposed an RF regression pipeline to correct baseline drift and environmental interferences in NO2 , PM10 , and PM2.5 measurements. The sensor drift and offset were handled by the adaptive iteratively reweighted penalised least squares (airPLS) techniques. However, despite its strengths, the RF method is inherently static and lacks memory of prior observations, which limits its capacity to generalize, model temporal drift, and long-term dynamics in sensor behaviour. Feature engineering is another important part of a calibration model. Having wellchosen features is essential for any sensor calibration pipeline, as the model can only learn relationships that are represented in its inputs. Prior work shows that explicitly representing period and trend noticeably improves spatio–temporal modelling in urban settings [20]. Modelling daily and weekly cycles enables learning recurrent pollutant dynamics, which are characteristic of urban air quality systems [21]. Deep learning methods that incorporate temporal structure, most notably Long ShortTerm Memory (LSTM) networks, have recently gained traction for their ability to model sequence-dependent patterns in sensor drift and performance [11]. Recurrent neural networks (RNNs) often suffer from vanishing or exploding gradients, which limit their ability to capture long-range dependencies in sequential data. LSTM networks address this issue

3

by using gated units that regulate how information is stored, updated, and discarded [22]. Temporal sequences calibrated with an LSTM network are well-suited to learning delayed effects and persistence in sensor signals [23]. These models view calibration not merely as static regression, but as a temporal inference task that evolves in response to environmental variability and site-specific dynamics. For example, Park et al. [11] developed a hybrid LSTM model for PM2.5 calibration, combining meteorological inputs such as temperature and relative humidity, and demonstrated superior capture of pollutant dynamics compared to feed-forward models. Veiga et al. [9] extended this by applying LSTM (alongside CNNs) in blind wireless sensor network calibration, showing marked error reductions without reference data. Beyond particulate matter, other studies highlight LSTM’s efficacy in calibrating gaseous pollutants. Han et al. [24] demonstrated that LSTM significantly improved NO2 accuracy, achieving R2 greater than 0.7 for the test set, surpassing RF and linear models in field conditions. Ryu and Park [25] introduced a band-sensitive LSTM that dynamically reweights errors in critical concentration ranges, reducing RMSE by 12% and improving tail accuracy. These examples underscore the flexibility of LSTM architectures for diverse pollutants and calibration scenarios. Meta-learning strategies are also garnering attention. Yadav et al. [26] proposed a few-shot calibration setup using Model-Agnostic Meta-Learning (MAML), enabling rapid calibration of new sensors with minimal co-location data. This paradigm addresses one of the central bottlenecks of ML-based calibration, which is the dependency on extensive labelled training data. Furthermore, Di Antonio et al. [27] focused on correcting PM measurements distorted by relative humidity. By incorporating a physical correction based on particle hygroscopicity, they drastically reduced PM2.5 overestimation. Alternatively, Patel et al. [28] proposed a hygroscopic-growth calibration for low-cost nephelometric PM2.5 sensors, modelling the seasonal evolution of kappa(κ) to correct RH-dependent biases. Additional studies further underscore the heterogeneity in LCS performance. Atfeh et al. [29] performed a comparative assessment of PM2.5 sensors under real-world conditions in Central Europe, noting that while many devices correlated well with reference monitors, bias and precision issues remained. They recommend hybrid techniques combining empirical correction and ensemble ML to improve accuracy. Other promising directions include AI-IoT integration. Rakib et al. [30] developed an IoT-based framework for air quality monitoring and 24-hour PM2.5 forecasting using an ARIMA model, demonstrating the potential scalability of cloud-connected sensor networks in real-time applications. The above studies demonstrate significant progress in sensor calibration but also reveal persistent challenges. The proposed work builds and improves upon the calibration framework developed by Bush et al. [19]. The calibration framework proposed by Bush et al. [19] relies on RF regression, which treats each observation independently and therefore cannot explicitly capture temporal dependencies, sensor drift, or delayed environmental

4

effects that influence low-cost sensor behaviour. These affect generalization to unseen data and have motivated the present work, which introduces a temporal deep learning calibration framework targeting PM2.5 , PM10 , and NO2 measurements. Several recent studies have shown that sequence-based deep learning models such as Long Short-Term Memory (LSTM) networks can improve calibration accuracy over traditional ML methods by modelling temporal dynamics in environmental sensor data [11, 24, 31]. Moving from point-wise prediction to sequence modelling allows the calibration framework to incorporate temporal context, enabling the model to learn persistence, delayed environmental effects, and gradual sensor drift that cannot be captured when each observation is treated independently. A hypertuned LSTM model has been developed and trained on co-located datasets, integrating lagged pollutant values and meteorological parameters. By incorporating rolling temporal windows and lagged environmental and pollutant features, the proposed method captures temporal dynamics in sensor responses and improves generalization performance on unseen datasets. The remainder of this paper is structured into four sections. Section 2 presents the methodology, covering data preparation, feature engineering, model architecture, and training procedures. Section 3 outlines the test case, and Section 4 reports the results, with performance metrics, comparisons against baseline models, and detailed error analysis. Finally, Section 5 concludes the study by summarizing key findings and directions for future research.

2.

Methodology

This section describes the steps followed in developing and evaluating the proposed calibration framework. The reliability of a calibration model comes not just from the choice of the model or architecture, but from how the supporting elements are designed. The design must ensure that sensor data are represented in a way that preserves their temporal structure and captures the environmental conditions under which they operate. Incorporating physically meaningful variables such as temperature, humidity, or flow dynamics allows the model to learn relationships grounded in sensor behaviour rather than relying on purely statistical correlations. Equally important is the selection of network parameters. Proper tuning of model parameters like window size, learning rate, and batch size ensures that the model captures long-term dependencies while remaining stable and generalizable. With these considerations, a calibration pipeline has been developed that treats LCS signals as temporally evolving processes, shaped by both environmental conditions and intrinsic sensor behaviour. The methodology has been summarized in Fig. 1. This consists of the following steps: • (a) Raw Sensor Data Input: Time-stamped measurements from low-cost devices (PM2.5 , PM10 , NO2 ) are combined with meteorological variables and co-located 5

reference observations to form the initial dataset. • (b) Feature Engineering & Preprocessing: Lagged and harmonic signals, domain-specific interactions, and traffic-related effects such as rush-hour are incorporated in the model. • (c) LSTM Calibration Model: A hypertuned LSTM network processes fixedlength sequences of input data, capturing sensor drift and delayed pollutant responses. • (d) Calibrated Output: Error metrics are then employed to evaluate the calibrated sensor output.

Figure 1: Overview of the calibration methodology: (a) raw sensor data input, (b) feature engineering and preprocessing, (c) LSTM calibration model, and (d) calibrated output evaluation.

2.1

Data Structure

The dataset analysed in this study is derived from the OxAria project [32], a large-scale deployment of low-cost air quality sensors across the city of Oxford. The measurements include particulate matter (PM2.5 and PM10 ) and nitrogen dioxide (NO2 ), alongside meteorological parameters such as ambient temperature and relative humidity. Depending on the pollutant channel, additional auxiliary variables are recorded. For electrochemical NO2 sensing, these include working and auxiliary electrode voltages, which provide information on baseline stability and environmental sensitivity. For particulate matter channels, flow-related variables such as sample flow rate and time-of-flight metrics are included, characterising transport through the optical chamber. An important aspect of the dataset used in this work is the baseline correction applied before analysis. As mentioned in the previous section, low-cost sensors frequently exhibit signal drift, offset variability, and spurious fluctuations due to environmental influences such as humidity or temperature. To address these issues, Bush et al. [19] implemented a four-stage preprocessing routine. The procedure begins with empirical filtering to remove physically implausible values and transient artefacts. This is followed by the application

6

of airPLS, which captures and subtracts slow-varying offsets without suppressing shortterm pollution dynamics. A compensation step is then introduced to prevent small overcorrections and residual anomalies, such as negative concentrations introduced by the correction. Baseline correction removes long-term drift and offsets that would otherwise distort the true pollution signal. This pipeline produces corrected sensor time series that are substantially more stable and reliable, while retaining the variability necessary for calibration against reference monitors. Bush et al. [19] reported that after correction, mean absolute errors of the Praxis Urban sensors are reduced, greatly improving the quality of the input data for machine learning applications. In the present study, this corrected dataset was used directly as the foundation for the model development. By relying on this preprocessed data, the analysis is able to concentrate on the design and evaluation of the temporal calibration rather than the intricacies of baseline adjustment. Table 1 summarizes the input variables used in model training by pollutant type. Feature (Variable)

NO2

PM10

PM2.5

Sensed concentration / mass (conc) Working electrode voltage (wev) Auxiliary electrode voltage (aev) Sample flow rate (sfr) Sample time of flight (mtf) Temperature (tmp) Relative humidity (hmd)

✓ ✓ ✓ – – ✓ ✓

✓ – – ✓ ✓ ✓ ✓

✓ – – ✓ ✓ ✓ ✓

Table 1: Sensor variables used for model training and prediction by pollutant type (adapted from Bush et al. [19]). This pollutant-specific feature matrix forms the foundation for the DL model trained in this study, ensuring consistency with the RF benchmarks used in prior work. In the LSTM-based model, additional temporal features and lag structures are incorporated during preprocessing, as described in the following sections.

2.2

Feature Engineering and Model Design

The engineered features focus on constructing a comprehensive input representation that reflects sensor behaviour, environmental drivers, and temporal structure. The model design integrates these features into a sequence-based learning pipeline, which includes data partitioning, normalization, sequence generation, and hyperparameter tuning. 2.2.1

Feature Engineering

To capture temporal dynamics and environmental influences, a rich set of features was engineered from the baseline-corrected LCS dataset. These include both direct sensor

7

observations and derived variables (Tab. 2). Here, the final input matrix X ∈ RT ×F consists of F features observed over T time steps, grouped into the following categories: • Raw Sensor Signals: Direct outputs from the Praxis units, including particulate concentrations, sample flow rate, time-of-flight metrics, ambient temperature, relative humidity, etc. These signals carry pollutant information but also embed sensor drift and environmental cross-sensitivities that require calibration. • Temporal Gradient Features: To capture rapid fluctuations, first- and secondorder percentage changes are computed [19]: pc15 x (t) =

x(t) − x(t − 1) , x(t − 1)

pc30 x (t) =

x(t) − x(t − 2) x(t − 2)

(1)

where x(t) denotes the value of a signal at time t. These emphasize short-term variability that raw concentrations may not fully capture. • Temporal and Cyclical Encodings: Time-of-day, day-of-week, and rush-hour indicators are included to capture diurnal and weekly cycles. To maintain continuity across cycle boundaries, harmonic features are encoded as:  2πh(t) , hour (t) = sin 24   2πd(t) sin day (t) = sin , 7 sin



 2πh(t) hour (t) = cos 24   2πd(t) cos day (t) = cos 7 cos



(2)

with h(t) and d(t) denoting hour and weekday, respectively. Monthly and seasonal cycles are similarly encoded. These features help the model recognize recurring temporal patterns in air pollution and improve consistency during repetitive events such as rush hours or seasonal shifts. • Interaction Features: Interaction is computed between various features available in the dataset. For example, the ratio mtf /(tmp + ϵ) and the product sf r · hmd, where ϵ is a small constant to avoid division by zero. These interactions reflect the combined influence of meteorology and particle optics on sensor performance [33]. • Target Lags: Recent lagged values of the pollutant under calibration and their percentage changes are included to provide autoregressive context. This allows the model to account for persistence and short-term correlation in pollutant levels.

2.2.2

Sequence Construction and Hyperparameter Tuning

Air pollutant dynamics and LCS readings depend not only on instantaneous conditions but also on recent history. To model this dependency, a supervised sequence-to-one learning strategy is adopted using a fixed-length rolling window technique. Let xt ∈ RF denote 8

Category

Examples

Raw Sensor Values conc, sfr, mtf, tmp, hmd Lagged Changes pc15 tmp, pc30 mtf, pc30 hmd etc. Time Features hour, day, rush hour Cyclical Encodings hoursin , daycos , monthsin , seasoncos etc. Sensor Interactions mtf/tmp, sfr·hmd etc. Table 2: Summary of raw and engineered features included in the calibration pipeline.

the engineered feature vector at time t (all inputs except the target channel), and let y t ∈ R be the co-located reference value used as the label. The dataset is then converted into overlapping input–output pairs as   X i = xi , xi+1 , . . . , xi+W −1 ∈ RW ×F ,

y i = y i+W ,

i = 1, 2, . . . , N,

(3)

where W is the sequence length (window size), T the total number of time steps, and N = T −W the total number of training pairs. This rolling-window setup preserves recent temporal context while also providing a fixed input size [34, 35]. Each window of W past measurements is mapped to the next-step target, and the window then slides forward by one step to form the next training pair. It ensures that each time step contributes both as part of an input sequence and, one step later, as a label. The effectiveness of sequence models such as LSTMs depends strongly on their hyperparameters, which control both the structure of the network and the dynamics of training. To identify a stable and accurate configuration, a grid search was carried out over three key parameters, W , learning rate (η), and batch size. Grid search systematically evaluates all possible combinations of these parameters within predefined ranges and selects the configuration that yields the best validation performance [36].

2.3

LSTM-Based Calibration Model

In this application, the input is a normalized (Eq. A3), fixed-length rolling window (Eq. 3) of engineered features (Eqs. 1 and 2), and the target is the reference observation. The network is implemented as a combination of multiple layers. A single LSTM block, followed by a dense layer with Leaky-ReLU (α = 0.01), a Dropout layer, and a final linear unit (Tab. 3) for the calibrated concentration. The model is trained with Adam using the mean absolute error objective with early stopping on validation loss. n

MAE =

1X |yi − ŷi | n i=1

(4)

where yi represents the true values, ŷi denotes the predicted values, and n is the total number of steps. 9

Layer Layer Details

Units / Params

Activation / Output Dim.

0 1 2 3 4

– 128 64 rate = 0.3 –

(W, F ) ∈ R128 Leaky-ReLU (α = 0.01); ∈ R64 ∈ R64 Linear; ŷt+1 ∈ R

Input LSTM Dense Dropout Dense (output)

Table 3: Layer specification for the LSTM calibration network (matching the implemented model). With such rich engineered inputs, the network has sufficient capacity to overfit these idiosyncrasies. L2 regularization (λ = 10−4 ) discourages large weights, which helps smooth the function and makes it generalize more reliably. Dropout prevents over-reliance on individual neurons, effectively acting like an ensemble to reduce variance [37]. Together with early stopping, these controls produce a stable calibrator that maintains accuracy when applied to hold-out and unseen periods.

2.4

Evaluation Metrics

The performance of the calibration strategy is assessed using three widely adopted statistical metrics: the coefficient of determination (R2 ), the Mean Absolute Error (MAE), and the Root Mean Squared Error (RMSE). These metrics capture different aspects of calibration accuracy. • R2 (Coefficient of Determination): R2 values quantify the proportion of variance in the reference observations explained by the predictions: 2

PN

R = 1 − Pi=1 N

(yi − ŷi )2

i=1 (yi − ȳ)

2

(5)

where yi are reference values, ŷi are the calibrated values, and ȳ is the mean of the reference set. A value close to 1 indicates strong explanatory power [38]. • MAE (Mean Absolute Error): MAE represents the average absolute deviation between predicted and observed values (Eq. 4). MAE provides a straightforward measure of typical prediction error and is less influenced by extreme outliers [38]. • RMSE (Root Mean Squared Error): RMSE emphasizes the larger deviations by squaring residuals before averaging: v u N u1 X t RMSE = (yi − ŷi )2 N i=1

(6)

RMSE is particularly relevant for air quality studies, as it penalizes missed peaks in pollutant concentrations [38]. 10

Together, these three metrics provide a balanced evaluation of calibration performance. R (Eq. 5) captures explanatory power, MAE (Eq. 4) and RMSE (Eq. 6) quantify error magnitudes under different sensitivities. 2

3.

Test Case

The OxAria LCS network deployed 16 low-cost sensor units across diverse urban microenvironments, with the goal of enabling high-resolution spatio-temporal air quality monitoring. This network provides coverage of key urban pollutants, complementing the sparse but highly accurate reference stations maintained under the UK’s Automatic Urban and Rural Network (AURN) programme. Each LCS unit is based on the Praxis Urban sensing platform from South Coast Science Ltd, integrating an Alphasense NO2 -A43F electrochemical sensor and an Alphasense N3 optical particle counter (OPC). The present study focuses on a single device co-located with the reference-grade AURN reference site at Oxford St. Ebbe’s (UKA00518). Reference data span different periods for each pollutant, May 2020 to September 2021 for PM10 and PM2.5, and May 2020 to May 2021 for NO2 . The co-location ensures that both the LCS device and the reference station are exposed to identical atmospheric conditions, enabling reliable supervised calibration. The setup has been summarized in Tab. 4. Figure 2 illustrates the temporal evolution of temperature and relative humidity during the observation period, demonstrating the variability that low-cost sensors are exposed to in the field. Figure 3 shows the time series for PM2.5 , PM10 , and NO2 , comparing AURN reference values with the sensor signals. Parameter

Value

Sensor platform Praxis Urban (South Coast Science Ltd.) Pollutants measured NO2 , PM10 , PM2.5 Sensor types Electrochemical (NO2 ), OPC (PM) Reference station Oxford St. Ebbe’s (AURN, UKA00518) Reference instruments Teledyne T200 (NO2 ), Palas FIDAS 200 (PM) Data resolution 15-minute averages Table 4: Summary of the OxAria test case setup used for calibration experiments.

4.

Results

This section presents the results of the proposed LSTM-based calibration approach for the low-cost air quality sensors. First, the scatter plots are presented, followed by the results of the evaluation metrics and the comparison of temporal predictions. The tuned hyperparameters used in the final model are summarized in A.2.

11

(a) Relative humidity

(b) Temperature

Figure 2: Time series of meteorological variables: (a) Relative humidity and (b) temperature.

4.1

Pollutant-wise Performance Evaluation

Since the behaviour and error characteristics of LCS vary across pollutants, the calibrated performance has been analyzed separately for PM2.5 , PM10 , and NO2 . This pollutant-wise breakdown highlights how the LSTM setup adapts to distinct sensing challenges. For each case, the model scatter plots, calibration accuracy against reference instruments, and a comparison of error statistics are presented in the sections below. 4.1.1

Scatter Plots

Scatter plots for PM2.5 , PM10 , and NO2 (Fig. 4) show that the calibrated predictions follow the 1:1 line closely across all data splits. The red dashed line indicates the 1:1 reference (perfect fit), and the green line represents a linear regression fit [39] between the calibrated and reference values. For each pollutant, the points form a tight cluster around the 1:1 line, and the fitted slopes remain close to one, indicating very little bias. This behaviour is consistent across the sets, suggesting that the model maintains its accuracy. Taken together, the results show stable calibration performance for all three pollutants, with strong alignment between the reference and the fitted regression across splits. Table 5 summarises the error metrics for all three pollutants. PM2.5 shows the strongest performance, with R2 values above 0.97 and low MAE and RMSE across all 12

(a) PM2.5 sensor values against AURN reference at Oxford St. Ebbe’s.

(b) PM10 sensor values against AURN reference at Oxford St. Ebbe’s.

(c) NO2 sensor values against AURN reference at Oxford St. Ebbe’s.

Figure 3: Comparison of low-cost sensor measurements with AURN reference data at Oxford St. Ebbe’s for (a) PM2.5 , (b) PM10 , and (c) NO2 . splits. The calibration performs well for PM10 , though generalization is slightly weaker than PM2.5 . The model still captures most of the variance, and the error levels between validation and test sets remain close, which indicates stable behaviour. For NO2 , the R2 values stay above 0.88, and both MAE and RMSE remain within a narrow range, show13

(a) PM2.5 (Validation)

(b) PM2.5 (Test)

(c) PM10 (Validation)

(d) PM10 (Test)

(e) NO2 (Validation)

(f) NO2 (Test)

Figure 4: Calibrated vs. reference concentrations across validation and test datasets for PM2.5 , PM10 , and NO2 . ing that the model tracks the main concentration patterns despite the higher variability typical of NO2 sensors.

14

Pollutant

Split

R2

MAE (µg/m3 )

RMSE (µg/m3 )

PM2.5

Train Validation Test

0.98 0.98 0.97

0.81 0.97 0.97

1.27 1.47 1.45

PM10

Train Validation Test

0.97 0.93 0.91

0.86 1.14 1.16

1.24 1.80 2.07

NO2

Train Validation Test

0.93 0.89 0.88

0.98 1.17 1.16

1.48 1.78 1.80

Table 5: Calibration performance metrics for PM2.5 , PM10 , and NO2 across training, validation, and test dataset splits.

4.2

Generalization to Unseen Data

To assess the performance of the trained LSTM model in unseen situations, model performance was evaluated on an unseen dataset collected at Oxford St. Ebbe’s between 23 and 30 September, 2021, for PM2.5 and PM10 , and between 23 and 30 May, 2021, for NO2 . This evaluation provides a practical check on model robustness, since real-world deployment often involves changes in weather, local activity, and instrument drift. 4.2.1

Results for unseen data

For PM2.5 , the unseen-data evaluation shows that the model follows the main variations of the reference series at both temporal resolutions (Fig. 5). At 15-minute resolution, the largest peaks tend to be slightly overestimated, although overall agreement remains strong. After averaging to hourly means, the predictions become less noisy and align more closely with the reference data. As summarized in Tab. 6, performance improves slightly at hourly resolution, with R2 increasing from 0.75 to 0.77, while MAE decreases from 1.63 to 1.56 µg/m3 and RMSE from 2.37 to 2.25 µg/m3 , indicating that the model benefits from reduced short-term variability. Type

R2

MAE (µg/m3 )

RMSE (µg/m3 )

1.63 1.56

2.37 2.25

Unseen Test (15-min) 0.75 Unseen Test (1H avg) 0.77

Table 6: PM2.5 performance on unseen data at 15-minute and hourly resolutions. The same conclusions can be drawn from the PM10 evaluation, which shows strong agreement between calibrated and reference concentrations at both temporal resolutions, although some peak values are slightly misrepresented (Fig. 6). After averaging to an hourly resolution (Tab. 7), short-lived fluctuations are smoothed, allowing a clearer comparison with the reference signal. Overall, the results indicate that the model maintains 15

(a) 15-minute resolution

(b) Hourly resolution

Figure 5: Time series of PM2.5 for unseen data: (a) reference and calibrated series at 15-minute resolution and (b) hourly-averaged reference and calibrated series. good performance under unseen conditions, with moderate increases in error compared to the testing period. Type

R2

MAE (µg/m3 )

RMSE (µg/m3 )

2.08 1.90

2.77 2.53

Unseen Test (15-min) 0.71 Unseen Test (1H avg) 0.74

Table 7: PM10 performance on unseen data at 15-minute and hourly resolutions. The unseen NO2 evaluation shows that the model tracks the general temporal behaviour of the reference series, including the dominant daily cycles, but struggles with rapid fluctuations and sharp peaks (Fig. 7). This behaviour is expected, as NO2 concentrations respond strongly to traffic emissions and meteorological variability, making short-term prediction more challenging. At 15-minute resolution, the model captures the main temporal structure but exhibits lower accuracy compared to the particulate pollutants. On the other hand, as summarized in Tab. 8, hourly averaging leads to improved performance, with R2 increasing by approximately 7.7%, while MAE and RMSE decrease 16

(a) 15-minute resolution

(b) Hourly resolution

Figure 6: Time series of PM10 for unseen data: (a) reference and calibrated series at 15-minute resolution and (b) hourly-averaged reference and calibrated series. by about 14.0% and 16.1%, respectively. This represents the largest relative improvement among the three pollutants and highlights the benefit of using longer periods for gases with higher short-term variability. Dataset

R2

MAE (µg/m3 )

RMSE (µg/m3 )

2.21 1.9

2.86 2.44

Unseen Test (15-min) 0.63 Unseen Test (1H avg) 0.71

Table 8: NO2 performance on unseen data at 15-minute and hourly resolutions. In summary, the unseen-data evaluations show that the LSTM calibration models perform consistently across pollutants and time scales. Although errors naturally increase under new conditions, especially for NO2 , the models still reproduce the main temporal behaviour and provide reliable calibrated outputs suitable for practical air-quality analysis. The results demonstrate that the LSTM framework can calibrate low-cost air quality sensor observations for all three pollutants to a high degree of accuracy. The framework also demonstrated strong performance on unseen data, confirming its effectiveness. The 17

(a) 15-minute resolution

(b) Hourly resolution

Figure 7: Time series of NO2 for unseen data: (a) reference and calibrated series at 15minute resolution and (b) hourly-averaged reference and calibrated series. calibration pipeline achieved consistently improved R2 values across training, validation, and test sets compared to the RF method reported by Bush et al. [19], while maintaining low error metrics. Further, the results have also been assessed using the Equivalence Spreadsheet Tool 3.1 (EC Working Group, 2020) [40], and the corresponding outputs are presented in Appendix A.1. In addition to the regression statistics, the tool provides estimates of uncertainty, which allows the calibrated sensor performance to be assessed against regulatory equivalence criteria. Application of the tool clearly shows that, when calibrated with the LSTM model, the low-cost sensor data can fully comply with the data quality objectives mandated by European regulation for objective estimation of air quality.

5.

Conclusion

This study presented a deep learning-based calibration framework for low-cost air quality sensors measuring PM2.5 , PM10 , and NO2 . The proposed approach combines a hyper18

tuned Long Short-Term Memory (LSTM) network with a rolling-window input structure, enabling the model to capture temporal dependencies and non-linear interactions between sensor signals and environmental variables. In contrast to traditional machine-learning approaches such as RF regression, which treat observations independently, the sequencebased LSTM architecture exploits the temporal structure of air-quality data. This allows the model to better represent persistence in pollutant concentrations, delayed environmental effects, and evolving sensor behaviour. Evaluation of co-located urban monitoring data demonstrated strong calibration performance across all pollutants. On test datasets, the calibrated outputs achieved R2 values exceeding 0.88 with mean absolute errors below 2 µg/m3 (or ppb for NO2), indicating a close agreement with reference-grade measurements. Importantly, the model maintained reliable performance on unseen temporal periods, confirming its ability to generalize beyond the training conditions. While particulate pollutants (PM2.5 and PM10 ) exhibited consistently high accuracy, NO2 calibration proved more challenging. Even so, the LSTM framework presented has demonstrated practical utility in delivering calibrated air quality data from low-cost sensors, which conform to the data quality objectives mandated by European air quality regulations. The results highlight the potential of sequence-based deep learning models for improving the reliability of low-cost sensor networks. By explicitly modelling temporal dynamics, the proposed framework addresses a key limitation of many existing calibration approaches based on static regression models, including RF methods commonly used in previous studies. This improvement is particularly relevant for urban monitoring applications where pollutant levels evolve rapidly, and sensor drift can affect long-term measurements. Several directions remain for future research. Extending the framework to multiple sensor deployments and different urban environments would allow assessment of spatial transferability, potentially supported by transfer learning or domain adaptation techniques. Incorporating uncertainty quantification through Bayesian neural networks, probabilistic forecasting, or ensemble approaches could provide confidence intervals for calibrated concentrations, which is important for regulatory and policy applications. In addition, integrating the calibration model into real-time edge computing systems could enable on-device processing and continuous recalibration, reducing latency in operational monitoring networks. Finally, exploring hybrid architectures, such as attention-based sequence models or temporal convolutional networks, may further enhance the ability to capture complex pollutant dynamics and improve robustness under changing environmental conditions.

Conflicts of Interest The authors declare no conflict of interest.

19

Code Availability The code developed for this study is available at: https://modelflows.github.io/ modelflowsapp/.

Acknowledgments The authors acknowledge the MODELAIR project that has received funding from the European Union’s Horizon Europe research and innovation programme under the Marie Sklodowska-Curie grant agreement No. 101072559. The results of this publication reflect only the author’s view and do not necessarily reflect those of the European Union. The European Union can not be held responsible for them. The authors gratefully acknowledge the Universidad Politécnica de Madrid (www.upm.es) for providing computing resources and Air Quality Consultants (www.aqconsultants.co.uk) for providing the dataset.

A.

Appendix

This Appendix A is organized into two sections. The first section presents the assessment of the LSTM Calibration model using the Equivalence Spreadsheet Tool 3.1, and the second section presents additional details on the DL model and parameters.

A.1

Assessment of LSTM Calibration Performance using the Equivalence Spreadsheet

European air quality legislation requires methods used for regulatory compliance monitoring to employ reference methods or methods demonstrated as equivalent to them [41]. To facilitate this assessment, the Commission provides a standardized spreadsheet-based equivalence tool that implements the statistical procedures outlined in the European air quality directives [40]. This tool evaluates candidate measurement methods against reference monitors through regression analysis and calculates key performance indicators, including expanded uncertainty, which must remain within prescribed thresholds: ≤25% for NO2 and ≤50% for both PM10 and PM2.5 . A.1.1

Equivalence Tool Output and Regression Analysis

The LSTM-calibrated sensor data was evaluated using the Equivalence tool, all at 15minute temporal resolution. An important methodological consideration emerged during this assessment regarding the calculation of the coefficient of determination (R2 ). Two distinct formulations are applied in the LSTM model and the Equivalence Tool, and they yield different values for the same datasets. They are computed as:

20

1. Direct R2 calculation: The standard sklearn implementation is computed as [42]: 2

PN

R = 1 − Pi=1 N

(yi − ŷi )2

2 i=1 (yi − ȳ)

(A1)

which quantifies the proportional reduction in variance explained by predictions relative to the mean baseline. This metric directly assesses prediction accuracy. 2. Regression R2 calculation: The Equivalence Tool employs the regression as [39]: ypred = β0 + β1 yref

(A2)

where, ypred and yref denote the calibrated and reference concentrations respectively, β1 is the regression slope, and β0 is the intercept. This formulation measures the strength of the linear association between predictions and reference values. Table A1 presents both R2 metrics for the LSTM calibrations alongside the Equivalence Tool results. For all three pollutants, the regression slopes are close to unity, and the intercepts are close to zero, confirming the absence of systematic bias. There is nonetheless a discrepancy between the direct (Eq. A1) and regression (Eq. A2) values, which is largest for PM2.5 , moderate for NO2 , and negligible for PM10 . For PM2.5 , the model overestimates sharp concentration peaks and under-predicts the low background levels that follow, where even small absolute errors are large relative to the typically low PM2.5 signal. This disproportionately suppresses the direct R2 while the regression R2 remains high as the overall trend is still captured. For NO2 , the calibrated values introduce discrepancies that reduce absolute accuracy without distorting the shape of the curve. Pollutant Direct R2 Regression R2 PM2.5 (µg m−3 ) 0.75 0.91 PM10 (µg m−3 ) 0.70 0.71 NO2 (ppb) 0.63 0.74 2 Table A1: Comparison of R metrics between direct calculation and OLS regression approach for LSTM-calibrated sensor measurements against the Equivalence Tool. With slope and intercept corrections applied, as permitted by the EU assessment framework, the LSTM-calibrated measurements meet all mandated data quality objectives. The expanded uncertainties are 22.11% for NO2 , 12.42% for PM10 , and 9.1% for PM2.5 , all falling within the prescribed thresholds of ≤25% for NO2 and ≤50% for PM10 and PM2.5 .

A.2

Model Development Parameters

This section details the normalization procedure, data partitioning strategy, and hyperparameter configuration adopted for the LSTM calibration model. 21

A.2.1

Normalization

All features are standardized using z-score normalization: X norm =

X −µ σ

(A3)

where X represents the input tensor, µ and σ are the mean and the standard deviation, respectively. This ensures consistent scaling across heterogeneous features and improves model stability [43, 44]. A.2.2

Data Partition

Reliable model evaluation requires a proper partition of the available data into training, validation, and test sets. In most cases, the division of the dataset typically depends on the case under study and is divided into ratios such as 80-10-10 or 70-20-10, etc [45, 46]. The dataset was split into training, validation, and test sets following a 70-20-10 ratio. This ensures that the model has sufficient examples to learn while retaining a meaningful portion for evaluation. A.2.3

Hyperparameter Tuning

The grid search systematically evaluated all parameter combinations from predefined ranges, selecting the best-performing configuration according to validation R2 . The search identified an effective setup, summarized in Tab. A2. Hyperparameter

Search Range

Optimal Value

Window Size (W ) Learning Rate (η) Batch Size

{8, 12, 16, 20, 24} 10−3 to 10−5 {16, 32, 48, 64}

12 0.0001 64

Table A2: Hyperparameter search space and optimal values used for the final LSTM model.

These settings produced the highest validation R2 and were subsequently fixed for testing and analysis. Such tuning is critical in environmental calibration, where data are often noisy and subject to drift, making stability just as important as accuracy.

References [1] World Health Organization, “Ambient (outdoor) air pollution.” Fact Sheet, 2024. Available at: https://www.who.int/news-room/fact-sheets/detail/ ambient-(outdoor)-air-quality-and-health. Accessed: 2025.

22

[2] V. A. Southerland, M. Brauer, A. Mohegh, M. S. Hammer, A. Van Donkelaar, R. V. Martin, J. S. Apte, and S. C. Anenberg, “Global urban temporal trends in fine particulate matter (pm2· 5) and attributable health burdens: estimates from global datasets,” The Lancet Planetary Health, vol. 6, no. 2, pp. e139–e146, 2022. [3] F. Karagulian, M. Barbiere, A. Kotsev, L. Spinelle, M. Gerboles, F. Lagler, N. Redon, S. Crunaire, and A. Borowiak, “Review of the performance of low-cost sensors for air quality monitoring,” Atmosphere, vol. 10, no. 9, p. 506, 2019. [4] T.-B. Ottosen, “Perspectives on the calibration and validation of low-cost air quality sensors,” Environmental Science & Technology, vol. 55, no. 19, pp. 12773–12775, 2021. [5] L. Liang, “Calibrating low-cost sensors for ambient air monitoring: Techniques, trends, and challenges,” Environmental Research, vol. 197, p. 111163, 2021. [6] K. Yin, J. Gersey, and P. Zhang, “In-field calibration of low-cost sensors through xgboost and aggregate sensor data,” arXiv preprint arXiv:2506.15840, 2025. [7] N. V. S. R. Nalakurthi, I. Abimbola, T. Ahmed, I. Anton, K. Riaz, Q. Ibrahim, A. Banerjee, A. Tiwari, and S. Gharbia, “Challenges and opportunities in calibrating low-cost environmental sensors,” Sensors, vol. 24, no. 11, p. 3650, 2024. [8] N. Jourdan, S. Sen, E. J. Husom, E. Garcia-Ceja, T. Biegel, and J. Metternich, “On the reliability of machine learning applications in manufacturing environments,” arXiv preprint arXiv:2112.06986, 2021. [9] T. Veiga, E. Ljunggren, K. Bach, and S. Akselsen, “Blind calibration of air quality wireless sensor networks using deep neural networks,” in 2021 IEEE International Conference on Omni-Layer Intelligent Systems (COINS), pp. 1–6, IEEE, 2021. [10] M. Si, Y. Xiong, S. Du, and K. Du, “Evaluation and calibration of a low-cost particle sensor in ambient conditions using machine-learning methods,” Atmospheric Measurement Techniques, vol. 13, no. 4, pp. 1693–1707, 2020. [11] D. Park, G.-W. Yoo, S.-H. Park, and J.-H. Lee, “Assessment and calibration of a lowcost pm2. 5 sensor using machine learning (hybridlstm neural network): Feasibility study to build an air quality monitoring system,” Atmosphere, vol. 12, no. 10, p. 1306, 2021. [12] P. Nowack, L. Konstantinovskiy, H. Gardiner, and J. Cant, “Machine learning calibration of low-cost no 2 and pm 10 sensors: Non-linear algorithms and their impact on site transferability,” Atmospheric Measurement Techniques, vol. 14, no. 8, pp. 5637– 5655, 2021.

23

[13] N. Zimmerman, A. A. Presto, S. P. Kumar, J. Gu, A. Hauryliuk, E. S. Robinson, A. L. Robinson, et al., “A machine learning calibration model using random forests to improve sensor performance for lower-cost air quality monitoring,” Atmospheric Measurement Techniques, vol. 11, no. 1, pp. 291–313, 2018. [14] I. D. Apostolopoulos, G. Fouskas, and S. N. Pandis, “Field calibration of a low-cost air quality monitoring device in an urban background site using machine learning models,” Atmosphere, vol. 14, no. 2, p. 368, 2023. [15] L. Spinelle, M. Gerboles, M. G. Villani, M. Aleixandre, and F. Bonavitacola, “Field calibration of a cluster of low-cost commercially available sensors for air quality monitoring. part b: No, co and co2,” Sensors and Actuators B: Chemical, vol. 238, pp. 706–715, 2017. [16] J. Han, W. Zhang, H. Liu, and H. Xiong, “Machine learning for urban air quality analytics: A survey,” arXiv preprint arXiv:2310.09620, 2023. [17] S. Mahajan and D. Helbing, “Dynamic calibration of low-cost pm2. 5 sensors using trust-based consensus mechanisms,” npj Climate and Atmospheric Science, vol. 8, no. 1, p. 257, 2025. [18] A. Bigi, M. Mueller, S. K. Grange, G. Ghermandi, and C. Hueglin, “Performance of no, no 2 low cost sensors and three calibration approaches within a real world application,” Atmospheric Measurement Techniques, vol. 11, no. 6, pp. 3717–3735, 2018. [19] T. Bush, N. Papaioannou, F. Leach, F. D. Pope, A. Singh, G. N. Thomas, B. Stacey, and S. Bartington, “Machine learning techniques to improve the field performance of low-cost air quality sensors,” Atmospheric Measurement Techniques Discussions, vol. 2021, pp. 1–34, 2021. http://doi.org/10.5194/amt-2021-282. [20] J. Zhang, Y. Zheng, and D. Qi, “Deep spatio-temporal residual networks for citywide crowd flows prediction,” in Proceedings of the AAAI conference on artificial intelligence, vol. 31, 2017. [21] S. Du, T. Li, Y. Yang, and S.-J. Horng, “Deep air quality forecasting using hybrid deep learning framework,” IEEE Transactions on Knowledge and Data Engineering, vol. 33, no. 6, pp. 2412–2424, 2019. [22] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT Press, 2016. http: //www.deeplearningbook.org. Accessed: 2025. [23] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997.

24

[24] P. Han, H. Mei, D. Liu, N. Zeng, X. Tang, Y. Wang, and Y. Pan, “Calibrations of low-cost air pollution monitoring sensors for co, no2, o3, and so2,” Sensors, vol. 21, no. 1, p. 256, 2021. [25] J. Ryu and H. Park, “Band-sensitive calibration of low-cost pm2. 5 sensors by lstm model with dynamically weighted loss function,” Sustainability, vol. 14, no. 10, p. 6120, 2022. [26] K. Yadav, V. Arora, M. Kumar, S. N. Tripathi, V. M. Motghare, and K. A. Rajput, “Few-shot calibration of low-cost air pollution (pm _{2.5}) sensors using meta learning,” IEEE Sensors Letters, vol. 6, no. 5, pp. 1–4, 2022. [27] A. Di Antonio, O. A. Popoola, B. Ouyang, J. Saffell, and R. L. Jones, “Developing a relative humidity correction for low-cost sensors measuring ambient particulate matter,” Sensors, vol. 18, no. 9, p. 2790, 2018. [28] M. Y. Patel, P. F. Vannucci, J. Kim, W. M. Berelson, and R. C. Cohen, “Towards a hygroscopic growth calibration for low-cost pm 2.5 sensors,” Atmospheric Measurement Techniques, vol. 17, no. 3, pp. 1051–1060, 2024. [29] B. Atfeh, Z. Barcza, V. Groma, Á. V. Tordai, and R. Mészáros, “Performance assessment of low-and medium-cost pm2. 5 sensors in real-world conditions in central europe,” Atmosphere, vol. 16, no. 7, p. 796, 2025. [30] M. Rakib, S. Haq, M. I. Hossain, and T. Rahman, “Iot based air pollution monitoring & prediction system,” in 2022 International Conference on Innovations in Science, Engineering and Technology (ICISET), pp. 184–189, IEEE, 2022. [31] X. Li, L. Peng, Y. Hu, J. Shao, and T. Chi, “Deep learning architecture for air quality predictions,” Environmental Science and Pollution Research, vol. 23, no. 22, pp. 22408–22417, 2016. [32] A. Bush, N. Papaioannou, F. Leach, F. D. Pope, A. Singh, G. N. Thomas, B. Stacey, and S. Bartington, “Sensor based ambient air concentration data for nitrogen dioxide and particles in oxford, measured by the oxaria project 2020 to 2021,” 2022. http://ora.ox.ac.uk/objects/uuid: 66fbe8c1-4b63-4124-bf0d-a78cbc9e1408. Accessed: 2025. [33] N. Castell, F. R. Dauge, P. Schneider, M. Vogt, U. Lerner, B. Fishbain, D. Broday, and A. Bartonova, “Can commercial low-cost sensor platforms contribute to air quality monitoring and exposure estimates?,” Environment international, vol. 99, pp. 293– 302, 2017. [34] L. B. Amor, I. Lahyani, and M. Jmaiel, “Recursive and rolling windows for medical time series forecasting: a comparative study,” in 2016 IEEE Intl Conference on Computational Science and Engineering (CSE) and IEEE Intl Conference on Embedded 25

and Ubiquitous Computing (EUC) and 15th Intl Symposium on Distributed Computing and Applications for Business Engineering (DCABES), pp. 106–113, IEEE, 2016. [35] L. Li, F. Noorian, D. J. Moss, and P. H. Leong, “Rolling window time series prediction using mapreduce,” in Proceedings of the 2014 IEEE 15th international conference on information reuse and integration (IEEE IRI 2014), pp. 757–764, IEEE, 2014. [36] J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” The journal of machine learning research, vol. 13, no. 1, pp. 281–305, 2012. [37] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research, vol. 15, no. 1, pp. 1929–1958, 2014. [38] D. Suriano and M. Penza, “Assessment of the performance of a low-cost air quality monitor in an indoor environment through different calibration models,” Atmosphere, vol. 13, p. 567, 03 2022. [39] C. R. Harris, K. J. Millman, S. J. Van Der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, et al., “Array programming with numpy,” nature, vol. 585, no. 7825, pp. 357–362, 2020. [40] European Commission, “Guide to the demonstration of equivalence of ambient air monitoring methods.” European Commission Working Group on Guidance for the Demonstration of Equivalence, 2010. http://environment.ec.europa.eu/topics/ air/air-quality/assessment_en. Accessed: 2025. [41] European Commission, “Directive 2008/50/ec of the European Parliament and of the Council of 21 May 2008 on ambient air quality and cleaner air for Europe,” 2008. http://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX: 32008L0050. Accessed: 2025. [42] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, et al., “Scikit-learn: Machine learning in python,” the Journal of machine Learning research, vol. 12, pp. 2825– 2830, 2011. [43] A. Hetherington, A. Corrochano, R. Abadía-Heredia, E. Lazpita, E. Muñoz, P. Díaz, E. Maiora, M. López-Martín, and S. Le Clainche, “Modelflows-app: data-driven postprocessing and reduced order modelling tools,” Computer Physics Communications, vol. 301, p. 109217, 2024. [44] A. Corrochano, G. D’Alessio, A. Parente, and S. Le Clainche, “Hierarchical higherorder dynamic mode decomposition for clustering and feature selection,” Computers & Mathematics with Applications, vol. 158, pp. 36–45, 2024. 26

[45] S. Kumar, “Data splitting technique to fit any machine learning model,” Received from: www. towardsdatascience. com, 2020. [46] I. Muraina, “Ideal dataset splitting ratios in machine learning algorithms: General concerns for data scientists and data analysts,” in Proceedings of the 7th International Mardin Artuklu Scientific Research Conference, pp. 496–504, Mardin Artuklu University, 2022.

27

Record · ID 126536 · SHA-256 522a861a402ae5b4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.