TREA-Net: A Transferable Residual Epidemiological Adaptation Network for Dengue Incidence Forecasting Inesh Shukla1 , Madhurima Panja2 , Tanujit Chakraborty2,3 , Chittaranjan Hens1 1
International Institute of Information Technology Hyderabad, India 2 SAFIR, Sorbonne University Abu Dhabi, United Arab Emirates 3 Sorbonne Center for Artificial Intelligence, Sorbonne Université, Paris, France [email protected], [email protected], [email protected], [email protected]
arXiv:2607.26854v1 [cs.LG] 29 Jul 2026
Abstract Accurate multi-week dengue forecasting supports timely vector-control interventions, outbreak preparedness, and healthcare resource allocation. However, newly established surveillance systems often lack the historical data needed to train reliable neural forecasting models. Although pretrained time-series models offer promising zero-shot forecasts, their cross-domain training may not capture local epidemiological dynamics. We propose TREA-Net, a Transferable Residual Epidemiological Adaptation Network for dengue forecasting under limited data. TREA-Net augments neural forecasting backbones with projections from an Environmental TimeSeries Susceptible–Infected–Recovered model and learns a lightweight gated residual correction transferable from datarich to data-scarce regions. Its node-invariant design accommodates surveillance systems with different numbers of locations, while target adaptation requires learning only two global parameters. We transfer knowledge from long-running dengue surveillance in Colombia and Nicaragua to 8-weekahead forecasting in Mexico and Malaysia using only 78 or 104 weeks of target data. Across five neural backbones and ten transfer settings, TREA-Net improves the corresponding backbone in 9 out of 10 settings, with statistically significant gains. When integrated with TiRex (a foundation model for forecasting), it achieves the lowest mean absolute error across all target datasets. Conformal prediction further maintains empirical coverage while reducing 8-week predictioninterval width by 29.6% in Mexico. These results demonstrate TREA-Net’s potential as a lightweight and portable earlywarning framework for health agencies with limited surveillance data.
Code — anonymous.4open.science/r/TREA-Net-C8F8/
Introduction Global dengue incidence more than doubled between 1990 and 2021, reaching approximately 60 million cases annually and accounting for nearly 2 million disability-adjusted lifeyears worldwide (Zhang et al. 2024; World Health Organization 2023). This burden is expected to increase as climate change expands the geographic range and transmission season of Aedes aegypti, exposing regions with limited surveillance capacity to greater outbreak risk (Lee et al. 2023). Reliable multi-week forecasting is therefore increasingly important for vector-control planning, healthcare resource allocation, and public-health decision-making. Forecasts up to
Figure 1: Data-rich surveillance systems (Colombia and Nicaragua) contain long historical dengue records, whereas data-scarce systems (Mexico and Malaysia) provide only 78 or 104 weeks of observations. Differences in surveillance history, epidemiological dynamics, and the numbers of monitored regions make direct transfer difficult, motivating a transferable epidemiologically informed forecasting framework (8-week-ahead dengue forecasting under limited target data).
eight weeks ahead can provide sufficient lead time for intervention while remaining informative about near-term transmission dynamics (Chen et al. 2018; Johansson et al. 2019; Chakraborty, Chattopadhyay, and Ghosh 2019). Yet this task is particularly difficult for newly established surveillance systems, where short historical records limit the ability of data-driven models to learn stable temporal and spatial patterns for multivariate dengue incidence datasets. Existing approaches for epidemic forecasting can be organized into three broad paradigms. The first comprises mechanistic compartmental models, such as susceptible–infectious–recovered (SIR) models (Kermack and McKendrick 1927) and environmental time-series SIR (ETSIR) models (Wang et al. 2025). By explicitly representing transmission processes and incorporating epidemiological or environmental knowledge, these models can remain informative when observations are scarce. Their reliance on simplified structural assumptions, however, limits their abil-
ity to represent nonlinear interactions, behavioral responses, policy interventions, and other external influences that shape real-world outbreaks (Deng et al. 2020; Rodrı́guez et al. 2023). The second paradigm uses data-driven learning to infer temporal and spatial dependencies directly from surveillance records. Recurrent, convolutional, graph-based, and transformer architectures can model complex nonlinear epidemic patterns (Deng et al. 2020; Xie et al. 2022; Kamarthi et al. 2023; Panja et al. 2023), but typically require long training histories. Many spatiotemporal models are also tied to a fixed graph or number of spatial units, making transfer to surveillance systems with different administrative structures difficult without architectural modification and substantial retraining (Rodrı́guez et al. 2021). More recently, large-scale time-series foundation models have achieved strong zeroshot performance through pretraining on massive crossdomain datasets (Panja et al. 2025; Das et al. 2024; Ansari et al. 2024; Auer et al. 2025). Although they reduce the need for target-domain training, their forecasts are driven primarily by statistical patterns learned during pretraining and do not explicitly incorporate local epidemiological mechanisms. The third paradigm combines the flexibility of neural forecasting with the structural knowledge encoded by compartmental models. Recent hybrid approaches, including Epidemiologically-Informed Neural Networks (EINNs) (Rodrı́guez et al. 2023), Epidemic-Guided Deep Learning (EGDL) (Barman et al. 2025), EARTH (Wan et al. 2025), and DINN (Cao et al. 2026), incorporate epidemiological information through mechanistic features, latent disease dynamics, architectural constraints, or physics-informed objectives (Rodrı́guez, Adhikari, and Prakash 2022). Although these methods can improve forecasting accuracy and epidemiological consistency, they are typically trained end-toend within a single surveillance system. Consequently, the learned neural representations may become tightly coupled to local transmission parameters, spatial structure, and data availability, making them difficult to transfer to systems with different numbers of monitored regions or only short observational histories. Such coupling can also complicate optimization in data-scarce settings and limits straightforward integration with frozen, zero-shot time-series foundation models. This raises the central problem considered in this work (also see Figure 1 for an illustration): How can mechanistic epidemiological knowledge learned from data-rich source regions be transferred to a structurally different, data-scarce target surveillance system while preserving a pretrained forecasting backbone and requiring only minimal targetspecific adaptation? To address this gap, we introduce TREA-Net (Transferable Residual Epidemiological Adaptation Network), a modular transfer-learning framework for dengue forecasting under limited target data. TREA-Net combines ETSIRbased mechanistic guidance with predictions from any pretrained forecasting backbone through a lightweight residual adapter. The adapter is learned from data-rich surveillance systems and transferred to data-scarce targets without mod-
ifying the backbone or assuming a fixed number of monitored regions. This provides a practical approach to portable multi-week dengue early warning for health agencies with limited data and computational resources. Our main contributions are: • Transferable mechanistic adaptation. We formulate epidemiological guidance as a latent residual correction using ETSIR projections and pretrained forecasts, rather than embedding mechanistic equations in the training loss. The point-wise, N-invariant adapter contains no node-specific parameters or fixed spatial topology, allowing transfer across surveillance systems with different numbers of administrative units while requiring only two global target-adaptation parameters. • Multi-country low-data evaluation. We transfer TREA-Net from Colombia and Nicaragua to Mexico and Malaysia using only 78 or 104 weeks of target observations. Across five forecasting backbones and ten transfer settings spanning 7–33 subnational units, TREA-Net improves the corresponding backbone in nine settings. We further combine it with conformal prediction to provide calibrated uncertainty, with improvements assessed through statistical significance testing. • Open framework for public-health forecasting. We will release the TREA-Net implementation and evaluation pipeline to support reproducible assessment of userdefined forecasting models under realistic low-data transfer settings. By reducing data, retraining, and computational requirements, the framework can support earlier vector-control planning, healthcare resource allocation, and risk-informed decision-making in resourceconstrained regions.
Related Work and Problem Formulation Epidemic Forecasting and Mechanistic–Neural Models. Classical compartmental models, including SIR, SEIR, TSIR, and their environmental extensions, encode diseasetransmission mechanisms and relevant covariates but require setting-specific calibration and rely on simplified assumptions (Kermack and McKendrick 1927; Bjørnstad, Finkenstädt, and Grenfell 2002; Wang et al. 2025). Datadriven approaches instead learn complex spatial and temporal dependencies from surveillance records. Cola-GNN and EpiGNN model inter-regional transmission through learned graph structures (Deng et al. 2020; Xie et al. 2022), while PROFHiT imposes probabilistic coherence across hierarchical time series (Kamarthi et al. 2023). These models are generally trained and evaluated on a fixed set of regions and do not explicitly address transfer between surveillance systems with different spatial dimensions. Hybrid approaches seek to combine both paradigms. EINNs and DINNs introduce compartmental dynamics through physics-informed learning (Rodrı́guez et al. 2023; Cao et al. 2026), whereas PISID integrates an SIR module with region-specific spatial embeddings (Fujita and Akutsu 2025). CALI-Net transfers representations from a historical influenza forecaster to a COVID-affected setting but
requires target-specific adaptation (Rodrı́guez et al. 2021). Although these methods improve epidemiological consistency or cross-setting adaptation, their learned components remain coupled to particular diseases, spatial identities, or target-specific optimization. They therefore do not provide a backbone-agnostic mechanism for transferring a frozen epidemiological correction across surveillance systems with different numbers of regions or for augmenting zero-shot foundation models using limited target data.
Finkenstädt, and Grenfell 2002) by combining recent incidence with nonlinear effects of precipitation and temperature. Let It denote dengue incidence at week t, and define the variance-stabilized response yt = log (1 + It ). Suppressing the regional index for clarity, the ETSIR projection is " # L X E ybt = c0 + cI yt−ℓ + gP (Pt ) + gT (Tt ) yt−1 , (1)
Parameter-efficient adaptation and N -invariant time series. Feature-wise Linear Modulation (FiLM) conditions intermediate representations through learned scale and shift parameters (γ, β) (Perez et al. 2018), while PETSA adapts frozen time-series forecasters using parameter-efficient testtime updates (Medeiros et al. 2025). Related work on variable-dimensional time series includes CPiRi, which promotes channel-permutation-invariant transfer through interchannel attention and channel-shuffling regularization (Xu et al. 2026), and channel-independent architectures such as PatchTST, which apply shared temporal weights across channels (Nie et al. 2023). This principle also underlies pretrained time-series models including TimesFM, Chronos, and TiRex (Das et al. 2024; Ansari et al. 2024; Auer et al. 2025), whose zero-shot performance on epidemic data has recently been examined (Panja et al. 2025). These methods establish parameter-efficient adaptation and invariance to channel count, but do not explicitly combine mechanistic epidemic projections with forecasts from an independently trained backbone. TREA-Net builds on this shared-weight principle by learning a point-wise epidemiological correction that transfers across different numbers of regions and requires only a two-parameter target adapter.
where L is the historical lag length and
Problem Formulation. We consider cross-geographical dengue forecasting from data-rich source surveillance systems to a data-scarce target system. Let the source and target systems contain Ns and Nt subnational regions, respectively, where Ns and Nt need not be equal. For each region, the model receives a 26 -week window of dengue incidence and associated covariates. The target system provides only K ∈ {78, 104} weeks of historical observations, corresponding to approximately 1.5 or 2 years of surveillance. Given forecasts from a pretrained neural backbone and mechanistic projections from an ETSIR model, the objective is to predict dengue incidence for the next H = 8 weeks in every target region. We seek to learn an epidemiologically informed correction from the source systems that can be transferred to the target without retraining or modifying the forecasting backbone. The transferable component must therefore remain applicable across surveillance systems with different numbers of regions and require only minimal estimation from the limited target history.
Methodology ETSIR Mechanistic Prior To encode climate-sensitive dengue dynamics, we adopt the Environmental Time-Series SIR (ETSIR) model (Wang et al. 2025). ETSIR extends classical TSIR dynamics (Bjørnstad,
ℓ=1
gT (Tt ) = cT Tt + c′T (Tt − HT )+ , gP (Pt ) = cP Pt + c′P (Pt − HP )+ ,
(x)+ = max(x, 0).
The piecewise-linear functions allow the effects of temperature and precipitation to change beyond thresholds HT and HP , reflecting the non-monotone influence of weather on mosquito abundance and transmission. The lagged-incidence term summarizes recent epidemic activity, including transmission persistence and susceptibledepletion effects. All ETSIR parameters are estimated using training-period observations only. We use multiple-stepahead (MSA) estimation, which minimizes cumulative forecast error over a multi-week horizon rather than optimizing only one-step predictions. For parameter vector θ, the objective is θ̂ = arg min θ
MSA X HX
t
2 E yt+h − ybt+h|t (θ) .
h=1
At test time, future temperature and precipitation are replaced by week-specific climatological averages computed exclusively from the training period, preventing information leakage. The resulting eight-week ETSIR trajectory is not used as a standalone forecast; instead, it provides TREANet with a compact mechanistic prior that complements the prediction from the neural backbone.
N -Invariant Gated Residual Correction b B, Y b E ∈ RH×N denote the backAt forecast origin t, let Y t t bone (any forecasting framework) and ETSIR forecasts, respectively, where H is the forecast horizon and N is the number of regions. Their (h, n)-th entries are written as B E ybt+h|t,n and ybt+h|t,n ′ for h = 1, . . . , H and n = 1, . . . , N . After MinMax normalization using training-period statisB E tics, let ỹt+h|t,n and ỹt+h|t,n denote the normalized forecasts. For each horizon-region pair, we construct xt,h,n = h iT B E B E ỹt+h|t,n , ỹt+h|t,n , ỹt+h|t,n − ỹt+h|t,n ∈ R3 . The correction module estimates a gate gt,h,n and a residual adjustment ∆t,h,n : gt,h,n = σ wg⊤ ReLU (Wg xt,h,n + bg ) + bg , ⊤ ∆t,h,n = w∆ ReLU (W∆ xt,h,n + b∆ ) + b∆ , B ỹt+h|t,n = ỹt+h|t,n + gt,h,n ∆t,h,n .
Both multilayer perceptrons use a hidden dimension of 64. The gate gt,h,n ∈ (0, 1) controls the contribution of the
Figure 2: Overview of the proposed TREA-Net framework. ETSIR provides a mechanistic prior; a shared N - invariant gated residual module learns transferable epidemiological corrections from multiple source geographies, and only two target-specific adaptation parameters (γ, β) are estimated for deployment in a new surveillance system. residual correction, while ∆t,h,n determines its direction and magnitude. The corrected forecasts are subsequently mapped back to the original scale using the corresponding inverse transformations. The same functions are applied independently to every (h, n) pair, with no horizon-specific, node-specific, or topology-dependent parameters. The module is therefore invariant to N and can be transferred across surveillance systems containing different numbers of regions. Moreover, the ETSIR projection enters as an input feature rather than through a mechanistic loss term, preserving the forecasting backbone and isolating the transferable epidemiological correction.
Multi-Source Training Let Gsrc denote the set of source geographies, and let Dg contain the correction-training samples from source g. Because the source datasets differ substantially in size, particularly, Colombia contributes approximately 2.4 times as many training windows as Nicaragua; naively pooling all samples would bias learning toward the larger source. We therefore define Dmin , Dmin = min |Dg | , wg = g∈Gsrc |Dg | and train the correction parameters θ using the reweighted objective X X Lsrc (θ) = wg ℓi (θ), g∈Gsrc
i∈Dg
where ℓi denotes the forecasting loss for sample i. This weighting equalizes the aggregate contribution of each
source geography without discarding observations. For every source, training samples are generated from all five forecasting backbones and pooled before optimizing the shared correction module. No backbone identifier is supplied to the module; it observes only the normalized backbone forecast, the ETSIR projection, and their absolute discrepancy. Consequently, the learned correction is encouraged to capture relationships that generalize across both source geographies and backbone architectures.
Two-Scalar Target Adaptation During deployment in a target geography g ⋆ , both the forecasting backbone and the source-trained correction parameb remain frozen. We introduce two target-specific global ters θ scalars, γg⋆ and βg⋆ , which adjust the magnitude and offset of the transferred correction. For each forecast horizon h and target region n, T B ỹt+h|t,n = ỹt+h|t,n + γg⋆ gt,h,n ∆t,h,n + βg⋆ ,
where gt,h,n and ∆t,h,n are produced by the frozen gated residual module. The two adaptation parameters are estimated using forecasting windows constructed from the first K ∈ {78, 104} weeks of target observations: X b γ bg⋆ , βbg⋆ = arg min ℓi (γ, β; θ). γ,β
i∈Dgadapt ⋆
Because γg⋆ and βg⋆ are shared across all horizons and regions, target adaptation introduces only two trainable parameters, irrespective of the number of administrative units Ng∗ .
The corrected forecasts are then mapped back to the original incidence scale. As examined in Section , replacing the global adapter with node-specific scale and shift parameters introduces 2Ng⋆ parameters but provides no consistent improvement and is more susceptible to overfitting under limited target data.
Experiments Datasets We use weekly subnational dengue surveillance data from four countries. The source pool comprises Colombia, with NCol = 33 departments observed over 835 weeks, and Nicaragua, with NNic = 18 departments over 981 weeks. These records span approximately 16 and 19 years, respectively, and represent mature surveillance systems. The target datasets comprise Malaysia, with NMal = 15 states and 270 available weeks, and Mexico, with NMex = 22 states and 208 available weeks. To emulate newly established surveillance programs, target adaptation is restricted to the first K = 104 weeks of each target series. Weekly temperature and precipitation covariates, together with their 14-, 28-, and 42-day rolling summaries, are aligned with dengue incidence for each region. We report mean absolute error (MAE) on raw dengue case counts after applying the inverse MinMax transformation. MAE is used because its linear scale provides a direct interpretation in terms of weekly case-count error, making it suitable for downstream resource-allocation analysis discussed in the Real-World Impact section.
Forecasting Backbones We evaluate five forecasting backbones: LSTM, N-HiTS, TCN, PatchTST (Nie et al. 2023), and TiRex (Auer et al. 2025). The first four models use a 26-week input window and a hidden dimension of 64. For backbone-only baselines, these models are trained using only the available target history. To generate source samples for learning the transferable correction, separate backbone instances are trained on the source surveillance systems. TiRex is a pretrained timeseries foundation model and is applied in a zero-shot setting without source- or target-specific fine-tuning. Under fixed preprocessing and inference settings, its forecasts are deterministic; hence, its backbone-only MAE has zero variation across repeated runs. For all backbones, TREA-Net operates only on their forecasts and does not modify their internal architectures or parameters.
Target Adaptation and Evaluation Protocol For each target geography, we consider limited-history regimes of K ∈ {78, 104} weeks. The first K−20 weeks are used to estimate the two target-adaptation parameters, while the subsequent 20 weeks are used for validation and early stopping. With a 26-week input window and an 8-week forecast horizon, the two regimes yield 25 and 51 adaptationtraining windows, respectively. All observations after week K are held out for final testing and are not used for model selection, scaling, or adaptation. The target adapter is optimized using Adam. Learning rates, batch sizes, and early-
stopping patience are reported explicitly in the Supplementary Material.
Results and Comparisons Table 1 reports MAE for Malaysia and Mexico under K ∈ {78, 104} weeks of target history. Across the five backbones and two target geographies, TREA-Net yields statistically significant improvements in 9 of 10 backbone-geography comparisons according to the Wilcoxon signed-rank test at α = 0.05 (p values are < 0.05). The standalone ETSIR model produces substantially larger errors, particularly for Malaysia at K = 78 (MAE = 1334.9), illustrating the difficulty of calibrating a purely mechanistic model from short histories. ARIMA is comparatively competitive in Malaysia but performs less consistently in Mexico. Incorporating the ETSIR residual constraint into an MLP reduces MAE by 12.0% on average and markedly lowers variability across random seeds. For example, the standard deviation decreases from 7.9 to 2.0 in Malaysia and from 14.2 to 0.4 in Mexico at K = 78. Nevertheless, ETSIR-PINN does not outperform the strongest sequence-model baselines, suggesting that mechanistic regularization alone does not compensate for limited temporal expressivity. Among the forecasting backbones, zero-shot TiRex achieves the lowest backbone-only MAE in all four target settings. N-HiTS is the strongest target-trained model in Malaysia, whereas LSTM and TCN exhibit greater variation across seeds. Applying the same backbone-agnostic TREANet correction improves 18 of the 20 backbone-targethistory combinations. Improvements are larger in Mexico, reaching 16.8% for PatchTST at K = 78 and 14.8% for TiRex at K = 104, while gains in Malaysia range from 0.2% to 5.1%. TREA-Net combined with TiRex achieves the lowest MAE in every target setting, reducing error by 7.1% on average relative to the zero-shot TiRex baseline. The two exceptions are TREA-Net-N-HiTS in Malaysia at K = 104, where MAE increases by 0.4%, and TREA-Net-PatchTST in Mexico at K = 104, where the difference is negligible.
Ablation Study Table 2 compares four configurations using the TiRex backbone at K = 104, with each ablated variant modifying one component of TREA-Net. Removing the ETSIR feature increases MAE by 7.5% in Malaysia and 11.5% in Mexico, demonstrating that the mechanistic projection provides information beyond the backbone forecast. Training the gated correction only on the target data performs 5.1% worse in Malaysia and is comparable in Mexico, indicating that multi-source training improves transfer without requiring a larger target-specific model. Finally, replacing the two global adaptation parameters with 2N node-specific parameters produces no meaningful improvement in Malaysia and changes performance by less than 1% in Mexico, supporting the simpler two-scalar adapter.
Table 1: MAE on raw case counts for Malaysia and Mexico under K ∈ {78, 104} weeks of target history, averaged over five seeds. Bold and Underlined values denote the best and second-best results in each column, respectively. Avg. ∆% is the mean percentage reduction in MAE achieved by each TREA-Net variant relative to its paired backbone across the four target-history settings and is reported only for TREA-Net rows. All TREA-Net variants use a single backbone-agnostic correction module trained on reweighted Colombia and Nicaragua source data, with only the target-specific scalars (γ, β) estimated locally. †TiRex is applied zero-shot and produces deterministic backbone forecasts; hence, its standard deviation is zero. Malaysia
Mexico
Tier
Model
K=78
K=104
K=78
K=104
Reference
ETSIR alone ARIMA
1334.9 75.7
557.2 83.4
72.0 89.0
66.7 96.7
Physics-informed
ETSIR-PINN
48.1
49.4
68.9
75.4
Backbone
LSTM NHiTS TCN PatchTST TiRex†
34.7 28.4 36.6 49.3 23.8
43.8 29.0 40.0 50.5 26.6
71.3 68.7 73.6 89.5 54.9
65.3 61.0 66.4 87.7 60.1
TREA-Net (Proposed)
TREA-Net-LSTM TREA-Net-NHiTS TREA-Net-TCN TREA-Net-PatchTST TREA-Net-TiRex
34.4 28.3 36.3 47.8 22.6
43.4 29.1 39.9 49.7 25.5
68.9 65.5 65.9 74.4 52.4
62.7 59.4 64.8 87.7 51.2
Table 2: Component ablation of TREA-Net with the TiRex backbone at K = 104, averaged over five seeds. Avg. ∆% denotes the mean percentage increase in MAE relative to TREA-Net-TiRex across Malaysia and Mexico; positive values indicate worse performance. In-domain correction trains the gated correction solely on the target data without source transfer or subsequent adaptation. Per-node adapter replaces the two global parameters (γ, β) with node-specific scaleshift pairs (γn , βn ), introducing 2N target parameters. Method
MAL
MEX
Avg ∆%
TREA-Net-TiRex (reference)
25.5
51.2
—
w/o ETSIR feature In-domain adapter Per-node (2N ) adapter
27.4 26.8 25.5
57.1 51.0 50.8
+9.5 +2.4 −0.4
Uncertainty Quantification via Conformal Prediction We construct 90% prediction intervals using Ensemble Batch Prediction Intervals (EnbPI) (Xu and Xie 2023), which updates calibration residuals sequentially and is therefore well suited to evolving epidemic conditions. Under the K = 104 regime, we compare zero-shot TiRex with TREANet-TiRex at nominal miscoverage level α = 0.10 (refer to Supplementary material). Figure 3 presents representative 8-week-ahead forecasts for four subnational units across Malaysia and Mexico. Each panel shows the observed series, the TiRex baseline, and the TREA-Net-TiRex forecast with its 90% conformal interval. In Negeri Sembilan and Morelos, TREA-Net corrects baseline trajectories that persist at overly high levels or continue increasing despite an observed decline. In Perak and Chia-
Avg ∆%
2.4 1.8 3.4 5.3 7.1
pas, where TiRex is already accurate, the correction remains small, and the intervals are comparatively narrow. Across the examples, the conformal bands cover nearly all future observations, consistent with the aggregate coverage results.
Gate Interpretability Figure 4 compares mean gate activation with relative ETSIR skill across subnational regions. In Malaysia, gate activation decreases as the ETSIR-to-persistence MAE ratio increases (ρ = −0.70, p = 0.004), indicating that TREANet attenuates the transferred correction where the mechanistic prior is less reliable. This spatial association emerges despite the gate being trained only through the forecasting objective, without access to the held-out ETSIR skill metric. No corresponding spatial relationship is observed in Mexico (ρ = 0.01, p = 0.954), where ETSIR performance is more homogeneous across states. Instead, gate activation varies primarily across forecast origins, with a significant negative association with prediction error (n = 97, ρ = −0.267, p = 0.008); this temporal association is negligible in Malaysia (n = 159, ρ = −0.015, p = 0.855). These results suggest that the gate adapts along the dominant source of predictive variation-spatially in Malaysia and temporally in Mexico.
Discussion: Real-World Impact and Limitations Dengue imposes substantial health and economic costs, exceeding 3 billion USD annually in Latin America and the Caribbean alone (Laserna et al. 2018). Eight-week forecasts can help health agencies anticipate changes in regional case burden and prioritize vector-control teams, diagnostic capacity, clinical staffing, bed availability, and medical supplies. Forecast improvements should not be inter-
Figure 3: Representative 8-week-ahead forecasts with 90% EnbPI prediction intervals under K=104. Black lines denote observed weekly cases, with the dotted vertical line marking the forecast origin. Blue lines show TREA-Net–TiRex forecasts and shaded regions their conformal intervals; red dashed lines show the uncorrected TiRex baseline. TREANet corrects substantial baseline deviations in Negeri Sembilan and Morelos while preserving accurate forecasts in Perak and Chiapas, with the observed trajectories contained within the prediction intervals.
preted directly as cases averted, since health outcomes depend on subsequent interventions and local operational constraints. Rather, lower case-count errors provide decisionmakers with earlier and more reliable information for allocating limited resources. TREA-Net is designed for surveillance systems that cannot train large forecasting models from scratch. Its transferable correction module contains approximately 5K parameters, while target adaptation estimates only two global scalars and completes on a CPU in under one minute in our implementation. This lightweight design may broaden access to mechanistically informed forecasting capabilities that would otherwise remain concentrated in data-rich and computationally well-resourced settings. Our analysis uses aggregate weekly case counts at the first subnational administrative level, obtained from public health sources and OpenDengue (Clarke et al. 2024), together with publicly available climate covariates. No patientlevel identifiers, addresses, or individual records are processed. Nevertheless, surveillance data may contain underreporting, changes in case definitions, and spatially uneven ascertainment. TREA-Net should therefore support, rather than replace, public-health judgment. Prospective deployment would require local validation, engagement with health authorities, appropriate institutional and data-governance review, and monitoring for unequal performance across regions. Several limitations remain. TREA-Net requires K ∈ {78, 104} weeks of target surveillance and sufficient climate information to fit a local ETSIR model; it does not address monthly reporting, deployment with no target history, or prolonged reporting disruptions. Future weather inputs are
Figure 4: Mean gate activation versus relative ETSIR skill across subnational regions at K = 104. The horizontal axis reports the ratio of ETSIR MAE to persistence MAE, with values above 1 indicating that ETSIR performs worse than persistence. Gate activation is significantly lower where ETSIR is less reliable in Malaysia (ρ = −0.70, p = 0.004), whereas no spatial association is observed in Mexico (ρ = 0.01, p = 0.954). currently approximated using training-period weekly climatology, which cannot anticipate anomalous conditions; operational meteorological forecasts could improve responsiveness to emerging climate-driven risk. Finally, evaluation is retrospective and limited to two target countries. Prospective studies are needed to determine whether improved forecasts lead to better resource allocation, outbreak preparedness, and health outcomes.
Conclusion We introduced TREA-Net, a lightweight transfer-learning framework for dengue forecasting under limited target data. By combining ETSIR projections with backbone forecasts through an N-invariant gated residual correction, TREA-Net transfers epidemiologically informed forecasting knowledge across surveillance systems without node-specific embeddings or backbone fine-tuning. Target adaptation requires estimating only two global parameters, enabling deployment across geographies with different numbers of administrative units. Across Mexico and Malaysia under K ∈ {78, 104} weeks of target history, TREA-Net consistently improves diverse neural and foundation-model backbones, with statistically significant gains in 9 of 10 backbone–geography comparisons. The framework also provides calibrated prediction intervals and requires only a small correction module, making it suitable for health agencies with limited surveillance data and computational capacity. Future work will extend TREA-Net to additional diseases and surveillance systems, incorporate operational weather forecasts instead of historical climatology, and investigate adaptation with shorter or disrupted target histories. Prospective evaluation with public-health partners is also needed to determine how forecast improvements translate into better intervention timing, resource allocation, and outbreak preparedness.
References Ansari, A. F.; Stella, L.; Turkmen, C.; Zhang, X.; Mercado, P.; Shen, H.; Shchur, O.; Rangapuram, S. S.; Pineda Arango,
S.; Kapoor, S.; Zschiegner, J.; Maddix, D. C.; Mahoney, M. W.; Torkkola, K.; Wilson, A. G.; Bohlke-Schneider, M.; and Wang, Y. 2024. Chronos: Learning the Language of Time Series. arXiv:2403.07815. Auer, A.; Podest, P.; Klotz, D.; Böck, S.; Klambauer, G.; and Hochreiter, S. 2025. TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context Learning. In Advances in Neural Information Processing Systems (NeurIPS). ArXiv:2505.23719. Barman, M.; Panja, M.; Mishra, N.; and Chakraborty, T. 2025. Epidemic-guided deep learning for spatiotemporal forecasting of tuberculosis outbreak. Machine Learning, 114(10): 213. Bjørnstad, O. N.; Finkenstädt, B. F.; and Grenfell, B. T. 2002. Dynamics of Measles Epidemics: Estimating Scaling of Transmission Rates Using a Time Series SIR Model. Ecological Monographs, 72(2): 169–184. Cao, J.; Cheng, Z.; Su, M.; and Hui, C. 2026. Inferring unobserved vector dynamics for dengue forecasting using physics-informed neural networks and mechanistic transmission models. BioRxiv preprint, concurrent work. doi:10.64898/2026.02.13.705693. Chakraborty, T.; Chattopadhyay, S.; and Ghosh, I. 2019. Forecasting dengue epidemics using a hybrid methodology. Physica A: Statistical Mechanics and its Applications, 527: 121266. Chen, Y.; Ong, J. H. Y.; Rajarethinam, J.; Yap, G.; Ng, L. C.; and Cook, A. R. 2018. Neighbourhood level real-time forecasting of dengue cases in tropical urban Singapore. BMC Medicine, 16(1): 129. Clarke, J.; et al. 2024. OpenDengue: a scalable global database of dengue case data. https://opendengue.org. Das, A.; Kong, W.; Sen, R.; and Zhou, Y. 2024. A DecoderOnly Foundation Model for Time-Series Forecasting. In Proceedings of the 41st International Conference on Machine Learning (ICML). Deng, S.; Wang, S.; Rangwala, H.; Wang, L.; and Ning, Y. 2020. Cola-GNN: Cross-Location Attention Based Graph Neural Networks for Long-Term ILI Prediction. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (CIKM). Fujita, S.; and Akutsu, T. 2025. Enhancing epidemic forecasting with a physics-informed spatial identity neural network. PLOS ONE. Funk, C.; Peterson, P.; Landsfeld, M.; Pedreros, D.; Verdin, J.; Shukla, S.; Husak, G.; Rowland, J.; Harrison, L.; Hoell, A.; and Michaelsen, J. 2015. The climate hazards infrared precipitation with stations—a new environmental record for monitoring extremes. Scientific Data, 2(1): 150066. Gorelick, N.; Hancher, M.; Dixon, M.; Ilyushchenko, S.; Thau, D.; and Moore, R. 2017. Google Earth Engine: Planetary-scale geospatial analysis for everyone. Remote Sensing of Environment, 202: 18–27. Johansson, M. A.; Apfeldorf, K. M.; Dobson, S.; Devita, J.; Buczak, A. L.; Baugher, B.; et al. 2019. An open challenge to advance probabilistic forecasting for dengue epidemics.
Proceedings of the National Academy of Sciences, 116(48): 24268–24274. Kamarthi, H.; Kong, L.; Rodrı́guez, A.; Zhang, C.; and Prakash, B. A. 2023. When Rigidity Hurts: Soft Consistency Regularization for Probabilistic Hierarchical Time Series Forecasting. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Applied Data Science Track). Kermack, W. O.; and McKendrick, A. G. 1927. A Contribution to the Mathematical Theory of Epidemics. Proceedings of the Royal Society A, 115(772): 700–721. Laserna, A.; Barahona-Correa, J. E.; Baquero, L.; Castañeda-Cardona, C.; and Rosselli, D. 2018. Economic impact of dengue fever in Latin America and the Caribbean: a systematic review. Revista Panamericana de Salud Pública, 42: e111. Lee, H.; Calvin, K.; Dasgupta, D.; Krinner, G.; Mukherji, A.; Thorne, P.; Trisos, C.; Romero, J.; Aldunce, P.; Barrett, K.; et al. 2023. Climate change 2023: synthesis report. Contribution of working groups I, II and III to the sixth assessment report of the intergovernmental panel on climate change. Medeiros, H. R.; Sharifi-Noghabi, H.; Oliveira, G. L.; and Irandoust, S. 2025. Accurate Parameter-Efficient Test-Time Adaptation for Time Series Forecasting. In Proceedings of the 42nd International Conference on Machine Learning (ICML). ArXiv:2506.23424. Mordecai, E. A.; Cohen, J. M.; Evans, M. V.; Gudapati, P.; Johnson, L. R.; Lippi, C. A.; Miazgowicz, K.; Murdock, C. C.; Rohr, J. R.; Ryan, S. J.; Savage, V.; Shocket, M. S.; Stewart Ibarra, A.; Thomas, M. B.; and Weikel, D. P. 2017. Detecting the impact of temperature on transmission of Zika, dengue, and chikungunya using mechanistic models. PLoS Neglected Tropical Diseases, 11(4): e0005568. Muñoz-Sabater, J.; Dutra, E.; Agustı́-Panareda, A.; Albergel, C.; Arduini, G.; Balsamo, G.; Boussetta, S.; Choulga, M.; Harrigan, S.; Hersbach, H.; et al. 2021. ERA5-Land: a state-of-the-art global reanalysis dataset for land applications. Earth System Science Data, 13(9): 4349–4383. Nie, Y.; Nguyen, N. H.; Sinthong, P.; and Kalagnanam, J. 2023. A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers. In Proceedings of the International Conference on Learning Representations (ICLR). Panja, M.; Chakraborty, T.; Nadim, S. S.; Ghosh, I.; Kumar, U.; and Liu, N. 2023. An ensemble neural network approach to forecast Dengue outbreak based on climatic condition. Chaos, Solitons & Fractals, 167: 113124. Panja, M.; Modak, O.; Younes, G.; and Chakraborty, T. 2025. Zero-Shot Forecasting of Epidemics. In NeurIPS 2025 Workshop on Backdoor Reasoning, Robustness, Trust and Safety in Time Series (BERT2S). OpenReview: h5ElB4guAd. Perez, E.; Strub, F.; de Vries, H.; Dumoulin, V.; and Courville, A. 2018. FiLM: Visual Reasoning with a General Conditioning Layer. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence.
Rodrı́guez, A.; Adhikari, B.; and Prakash, B. A. 2022. When, why and how to build a hybrid mechanistic and deep learning model for epidemiology. arXiv preprint arXiv:2211.08271. Rodrı́guez, A.; Cui, J.; Ramakrishnan, N.; Adhikari, B.; and Prakash, B. A. 2023. EINNs: Epidemiologically-Informed Neural Networks. In Proceedings of the 37th AAAI Conference on Artificial Intelligence (AI for Social Impact track). Rodrı́guez, A.; Muralidhar, N.; Adhikari, B.; Tabassum, A.; Ramakrishnan, N.; and Prakash, B. A. 2021. Steering a Historical Disease Forecasting Model Under a Pandemic: Case of Flu and COVID-19. In Proceedings of the 35th AAAI Conference on Artificial Intelligence. ArXiv:2009.11407. Wan, G.; Liu, Z.; Shan, X.; Lau, M. S.; Prakash, B. A.; and Jin, W. 2025. EARTH: Epidemiology-Aware Neural ODE with Continuous Disease Transmission Graph. In Fortysecond International Conference on Machine Learning. Wang; et al. 2025. An environmental time-series SIR model with multiple-step-ahead estimation for dengue forecasting. (as cited in the main paper). World Health Organization. 2023. Dengue and Severe Dengue: Global Situation. WHO Disease Outbreak News. Accessed November 2025. Xie, F.; Zhang, Z.; Li, L.; Zhou, B.; and Tan, Y. 2022. EpiGNN: Exploring Spatial Transmission with Graph Neural Network for Regional Epidemic Forecasting. arXiv preprint arXiv:2208.11517. Xu, C.; and Xie, Y. 2023. Conformal prediction for time series. IEEE transactions on pattern analysis and machine intelligence, 45(10): 11575–11587. Xu, J.; Zhang, W.; Jing, X.; Nie, J.; Chen, S.; and Zhang, S. 2026. CPiRi: Channel Permutation-Invariant Relational Interaction for Multivariate Time Series Forecasting. In Proceedings of the International Conference on Learning Representations (ICLR). ArXiv:2601.20318. Zhang, S.-X.; Yang, G.-B.; Zhang, R.-J.; Zheng, J.-X.; Yang, J.; Lv, S.; Duan, L.; Tian, L.-G.; Chen, M.-X.; Liu, Q.; et al. 2024. Global, regional, and national burden of dengue, 1990–2021: findings from the global burden of disease study 2021. Decoding Infection and Transmission, 2: 100021.
Environmental Prior: Motivation and Extraction Nonlinear Climate Dependence of Dengue Transmission Dengue transmission is mediated by Aedes aegypti, whose abundance and vectorial capacity respond to temperature and precipitation in nonlinear and often non-monotonic ways. This motivates using an environmentally informed mechanistic prior, rather than representing climate effects only through unconstrained linear covariates. Temperature. Transmission potential is typically unimodal in temperature because biting rate, viral extrinsic incubation, and mosquito survival are each optimized over an intermediate thermal range. Risk increases toward an
Figure 5: Empirical climate–dengue associations in the Mexico surveillance data (2017–2020). Daily climate series from Google Earth Engine are aggregated weekly and aligned with dengue incidence at a four-week lead. Case counts are normalized within each state so that the curves reflect response shape rather than differences in baseline burden. (a) Incidence increases with temperature to a broad optimum around 23–29◦ C and declines at higher temperatures. (b) Incidence rises with moderate rainfall and then saturates. Error bars denote ±1 SEM. optimum near 29◦ C, but declines at cooler and hotter extremes, becoming negligible below approximately ∼18◦ C and above ∼34◦ C (Mordecai et al. 2017). A single linear temperature effect cannot represent this reversal: the same increase in temperature may raise transmission below the optimum but reduce it above the optimum. Precipitation. Rainfall exerts competing effects on mosquito abundance (Panja et al. 2023). Moderate precipitation creates and replenishes standing-water habitats, whereas intense rainfall can flush larvae and eggs from breeding containers. Dengue risk may therefore increase with rainfall up to a threshold and then saturate or decline. The piecewise environmental response in ETSIR provides a parsimonious way to represent such threshold-dependent effects. Figure 5 illustrates these relationships in the target data. Dengue incidence exhibits a hump-shaped association with temperature and a saturating, non-monotonic association with rainfall. These patterns suggest that a single linear covariate effect is inadequate, whereas the breakpoint structure in ETSIR can represent changes in both the magnitude and direction of environmental influence. Why ETSIR? These biological relationships motivate the ETSIR prior used in the main paper. Instead of treating temperature and precipitation as globally linear predictors, ETSIR represents their effects through piecewise-linear forcing functions gT (·) and gP (·) with learned breakpoints HT and HP . The rectified terms ( Tt − HT )+ and (Pt − HP )+ allow the marginal effects of temperature and rainfall to change beyond these thresholds, thereby accommodating thermal optima, saturation, and potential wash-out effects. The resulting ETSIR projection provides TREA-Net with a compact climate-aware summary of dengue transmission dynamics. Consistent with the ablation results in the main paper, this mechanistic signal contributes information beyond that contained in the backbone forecast alone.
Climate Data Extraction The climate covariates used by ETSIR are obtained from publicly available satellite-derived and reanalysis products accessed through Google Earth Engine (Gorelick et al. 2017). A common extraction and aggregation pipeline is applied across all source and target geographies to ensure consistent and comparable environmental features. • Precipitation. Daily precipitation is obtained from the CHIRPS dataset (UCSB-CHG/CHIRPS/DAILY) (Funk et al. 2015), which provides quasi-global rainfall estimates at approximately ∼ 5.5 km spatial resolution. • Temperature. Daily 2m air temperature is obtained from ERA5-Land (ECMWF/ERA5 LAND/DAILY AGGR) (MuñozSabater et al. 2021) and converted from Kelvin to degrees Celsius. Missing pixels along coastal boundaries are filled using a 3 × 3-pixel focal mean before aggregation to the administrative-region level. • Vector-viability mask. Both climate products are restricted to elevations below 2000 m using the SRTM digital elevation model (USGS/SRTMGL1 003). Because Aedes aegypti occurrence and dengue transmission are substantially reduced at high elevations, this mask prevents cold, high-altitude pixels that are unlikely to support transmission from diluting the administrative-region climate summaries. • Spatial aggregation. The masked daily rasters are spatially averaged within each first-level administrative polygon obtained from FAO/GAUL/2015/level1. The resulting regional climate series are then matched to the department or state names used in the corresponding dengue surveillance data. • Temporal aggregation and lagged features. Daily climate series are aggregated to epidemiological weeks, and rolling 14-, 28-, and 42-day means are constructed to capture delayed environmental effects on mosquito development and dengue transmission. At operational inference time, climate covariates over the forecast horizon are replaced by week-specific climatological averages estimated exclusively from the training period, as described in the ETSIR Prior section of the main paper. The extraction pipeline is used solely for data preparation. It operates on publicly available aggregate climate rasters and first-level administrative boundaries and produces one regional table per geography with columns (date, department, precip mm, temp c). These tables are subsequently used as inputs to the ETSIR estimation procedure.
Additional Experimental Details Dataset Details Table 3 summarizes the four subnational dengue surveillance datasets. The source–target assignment reflects surveillance maturity: Colombia and Nicaragua each provide more than 15 years of weekly case records and constitute the transfer-source pool, whereas Mexico and Malaysia
Table 3: Subnational dengue surveillance datasets. “Units” denotes the number of first-level administrative units (departments or states) with usable dengue case and climate series. Target datasets are evaluated under K ∈ {78, 104} weeks of available local history. Country
Role
Units
Weeks
Colombia Nicaragua Mexico Malaysia
Source Source Target Target
33 18 22 15
835 981 208 270
Surveillance span ∼16 years ∼19 years 2017–2020 2010–2015
are used as targets to emulate newly established surveillance programs. For each target, only the first K ∈ {78, 104} weeks are available for local model training and adaptation, and all subsequent observations are reserved for out-ofsample testing. Each geography is paired with the CHIRPS and ERA5-Land climate covariates described in Appendix , including temperature, precipitation, and their 14-, 28-, and 42-day rolling summaries, aligned weekly for each administrative unit.
Horizon-Wise Prediction-Interval Efficiency Although empirical coverage remains broadly stable, TREA-Net changes the width of the conformal prediction intervals, particularly at longer forecast leads. In Mexico, the uncorrected TiRex backbone produces narrower intervals at short leads (h ∈ {1, 2}), consistent with its strong near-term autoregressive performance. As the forecast lead increases, however, the TiRex intervals widen more rapidly. At h = 8, TREA-Net–TiRex produces intervals that are 29.6% narrower than those of TiRex while maintaining comparable empirical coverage (Figure 6). This result suggests that the ETSIR-informed residual correction improves interval efficiency at longer leads by stabilizing the underlying point forecasts. In Malaysia, TREA-Net and TiRex produce intervals of comparable width across the forecast horizon, consistent with the smaller corrections observed in this target geography. Overall, these results indicate that incorporating the mechanistic projection can yield tighter prediction intervals at longer forecast leads without an evident loss of empirical coverage.
Empirical Coverage under Distribution Shift Both methods exhibit some under-coverage relative to the nominal 90% level, attaining approximately 87% empirical coverage in Mexico and 81% in Malaysia. This shortfall coincides with a shift from predominantly low-incidence validation periods to test periods containing larger epidemic peaks, for which forecast residuals are substantially greater. Despite this shift, TREA-Net preserves the empirical coverage of the corresponding zero-shot TiRex baseline. Thus, the ETSIR-informed correction improves point forecasts and interval efficiency without an observed deterioration in empirical coverage.
Figure 6: Average EnbPI prediction-interval width by forecast lead in Mexico under K = 104 and nominal miscoverage level α = 0.10. The shaded region indicates forecast leads for which TREA-Net–TiRex produces narrower intervals than the TiRex baseline. Both methods attain approximately 87% empirical coverage over the test period.
Statistical Significance Testing For each backbone–target pair, we compare TREA-Net with its corresponding backbone across the two target-history regimes, K ∈ {78, 104}, and five random seeds. This yields 10 paired MAE differences per backbone–target comparison. We apply a two-sided Wilcoxon signed-rank test to these paired differences at significance level α = 0.05. TREA-Net achieves statistically significant improvements in 9 of the 10 backbone–target comparisons.
Implementation Details and Compute Footprint Table 4 summarizes the architecture and optimization settings for all trainable components. The correction module and the trainable forecasting backbones are optimized with early stopping using a held-out validation window and a patience of 20 epochs. TiRex is applied zero-shot and is therefore not updated. At deployment, the backbone and the source-trained correction module remain frozen, while only the two target-specific parameters (γ, β) are estimated from the adaptation portion of the available local target history.
Compute Footprint The transferable component is the source-trained correction module, containing approximately 5,000 parameters, which is trained once on the source pool and subsequently frozen. For a new target geography, only the two global adaptation parameters (γ, β) are estimated. In our implementation, this optimization completes in under one minute on a CPU and does not require GPU acceleration. Backbone and ETSIR forecasts are generated once and cached before adaptation. Consequently, the incremental computation required to adapt TREA-Net to a new surveillance system is small relative to retraining or fine-tuning the forecasting backbone.
Table 4: Architecture and optimization settings for all model components. Maximum epoch counts are upper bounds; early stopping with a patience of 20 epochs typically terminates training earlier. TiRex is applied zero-shot and its parameters were not updated. Setting
Value
Input sequence length (IL) Forecast horizon (H) Validation window Seeds
26 weeks 8 weeks 20 weeks 1–5
Correction module (Gate+Delta, N -invariant) hidden dimension 64 dropout / gate-bias init 0.5 / 0.2 parameters ∼5K optimizer / learning rate Adam / 5×10−5 max epochs / patience 300 / 20 batch size 32 Neural backbones (LSTM, NHiTS, TCN, PatchTST) hidden dimension 64 optimizer / learning rate Adam / 10−3 max epochs / patience 300 / 20 TiRex backbone zero-shot (untrained) (γ, β) target adapter trainable parameters optimizer / learning rate max epochs / patience fit cost
2 scalars Adam / 10−2 200 / 20 < 1 min, CPU
Conformal intervals (EnbPI)
α = 0.10 (90% nominal)