Generated using the official AMS LATEX template v6.1 two-column layout. This work has been submitted for publication. Copyright in this work may be transferred without further notice, and this version may no longer be accessible.
arXiv:2605.16163v1 [physics.ao-ph] 15 May 2026
SwAIther-Precip: Lead-Time-Aware Bias Correction Enables Kilometer-Scale Downscaling of Global AI Precipitation Forecasts over Switzerland Dan Assoulinea , Erwan Kochb , Federico Amatoa , Filippo Quarenghib,c , Daniele Nerinid , Thibaut Loiseaua , Kyle van de Langemheen a , and Tom Beuclerb,c . a Swiss Data Science Center, ETH, Zürich, Switzerland b Expertise Center for Climate Extremes, University of Lausanne, Lausanne, Switzerland c Faculty of Geosciences and Environment, University of Lausanne, Lausanne, Switzerland d Federal Office of Meteorology and Climatology MeteoSwiss, Locarno-Monti, Switzerland
ABSTRACT: Skillful medium-range precipitation forecasting at kilometer scale remains challenging over complex terrain because precipitation arises from multiscale nonlinear processes that global models cannot explicitly resolve at affordable cost. Global AI weather models can produce skillful medium-range forecasts, but their native 0.25° resolution limits direct use for local hazard applications. Statistical downscaling can help bridge this gap, yet existing approaches often struggle with state-dependent, and especially lead-time-dependent, biases in global forecasts. We introduce SwAIther-Precip, a lead-time-aware downscaling framework that converts coarse-resolution AIFS forecasts into probabilistic km-scale precipitation fields over Switzerland. First, a U-Net conditioned on lead time via feature-wise linear modulation deterministically corrects systematic biases at coarse resolution. This targeted correction enables a cheaper super-resolution stage conditioned only on corrected precipitation, allowing direct training on observations rather than on the full atmospheric state. A diffusion-based model then generates fine-scale spatial variability independently of lead time. Using AIFS forecasts and CombiPrecip radar–gauge observations, SwAIther-Precip reduces CRPS by 48% relative to raw AIFS. The generated fields reproduce observed spatial variability with spectral fidelity above 0.85 at large scales and 0.88 at small scales, corresponding to an effective resolution of ≈4 km on a 1 km grid for lead times up to 5 days. Training across lead times further improves long-range performance, yielding a 13% CRPS reduction at 6 days relative to lead-time-specific models. These results show that explicitly correcting lead-time-dependent biases before generative super-resolution is key to efficient km-scale probabilistic downscaling of global AI precipitation forecasts.
SIGNIFICANCE STATEMENT: AI weather models can now make skillful global forecasts up to two weeks ahead, but they are still too coarse for local precipitation hazards in mountainous regions such as Switzerland, which is a stringent testbed because terrain strongly shapes precipitation and high-quality observations allow rigorous evaluation. We show that correcting forecast errors in a lead-time-aware way before adding fine-scale detail makes these models much more useful locally. Using forecasts from a public global artificial intelligence weather model and Swiss radar–rain gauge observations, our method produces kilometer-scale precipitation forecasts, cuts forecast error in two relative to the raw model, and preserves realistic spatial patterns out to 6 days.
by global forecasting systems. In conventional NWP, the physically relevant quantity is the effective spatial resolution, i.e., the smallest wavelength at which the model retains realistic variance before the power spectrum becomes distorted (Skamarock 2004), which is often several grid lengths coarser than the nominal mesh (Klaver et al. 2020). Reaching km-scale effective resolution means entering a non-hydrostatic regime tightly coupled with convection. In NWP, this forces a trade-off between horizontal resolution, ensemble size, and forecast range; AI can more readily target the spatiotemporal resolution of interest. Deep generative radar nowcasting predicts highresolution precipitation minutes ahead (Ravuri et al. 2021), while km-scale AI emulators such as StormCast (Pathak et al. 2026) and HRRRCast (Abdi et al. 2026) extend this capability to the first few forecast hours. However, kmscale AI weather prediction at medium range (2 days to 2 weeks) remains rare, in part because forecast errors accumulate rapidly when a km-scale atmospheric state is rolled out over many time steps.
1. Introduction Accurate precipitation forecasting at high spatial resolution remains a central challenge in numerical weather prediction (NWP), particularly over regions with complex topography such as Switzerland. Precipitation fields exhibit strong spatial heterogeneity and nonlinear dynamics driven by orography, mesoscale organization, and local atmospheric conditions, which are often poorly resolved
In parallel, global AI weather prediction has advanced rapidly. Pangu-Weather (Bi et al. 2023) and GraphCast (Lam et al. 2023) established that neural networks trained on meteorological reanalyses can produce skillful mediumrange deterministic global forecasts. These systems are now routinely evaluated with common benchmarks such as
Preprint version with supplemental material appended. Corresponding author: Dan Assouline, [email protected]
1
2 WeatherBench 2 (Rasp et al. 2024). In this work we choose AIFS, the global AI weather prediction system of the European Centre for Medium-Range Weather Forecasts, as the coarse driver because it remains deterministic, which keeps post-processing inexpensive, and because it provides the multivariate atmospheric context needed for precipitation correction (Lang et al. 2024). This choice is especially appropriate for precipitation because recent AIFS updates have both improved precipitation skill and expanded the set of output variables (Moldovan et al. 2025). Even so, current global AI forecasts remain coarse relative to the kilometer scale, typically around 25–30 km, which limits their direct use for local and regional decision-making. Bridging this spatial-resolution gap follows two broad strategies. The first learns or emulates regional atmospheric evolution directly. Machine-learning limited-area models (LAMs) have made this direction increasingly realistic (Adamov et al. 2025): CRPS-LAM trains regional probabilistic forecasting with proper scoring rules rather than only with diffusion-based samplers (Larsson et al. 2025), while other systems reach 3 km and 1 h resolution for non-precipitation surface variables (Xu et al. 2025). The design space of this first strategy extends beyond boundary-forced LAMs: stretched-grid models refine the mesh over the target region without explicit lateral boundaries (Nipen et al. 2025), and OneForecast combines global and regional prediction within a single multiscale framework (Gao et al. 2025). Hybrid AI–physics systems offer a lower-cost alternative; for example, Pangu-Weather can drive a regional WRF configuration targeting extreme precipitation (Xu et al. 2024), while Sha et al. (2026) learn a stable AI regional emulator from reanalysis and climatemodel forcings, and related work has sought to reduce the cost of dynamical downscaling through ML surrogates (Hobeichi et al. 2023). Despite its physical appeal, this first strategy inherits the main burdens of conventional regional modeling: lateral-boundary treatment, rollout stability, and the need to predict the full atmospheric state at high resolution. STCast, for example, explicitly addresses boundary mismatch through spatially aligned attention, underscoring the importance of such issues (Chen et al. 2025). The second strategy, which we adopt here, bypasses many of these limitations through forecast refinement: a completed coarse forecast is downscaled post hoc to local scales without evolving the regional atmospheric state at kilometer resolution. CNN-based post-processing has been shown to improve precipitation forecasts from weathermodel output (Badrinath et al. 2023), and multivariate approaches exploit the full set of NWP predictors (RojasCampos et al. 2023). Generative methods have progressed from GAN and VAE-GAN stochastic downscaling of imperfect IFS forecasts (Harris et al. 2022), to joint bias correction and downscaling via CorrectorGAN (Price and
Rasp 2022), and to generative post-processing for regional rainfall over East Africa (Antonio et al. 2025). CorrDiff introduced a coarse-regression-plus-residual-diffusion decomposition for km-scale atmospheric downscaling from coarse inputs (Mardani et al. 2025), a successful framework since extended to forecasting, for example to downscale forecasts over China at 3 km resolution for lead times up to 72 h (Sun et al. 2026). This recent progress motivates extending forecast refinement to longer lead times, especially for km-scale precipitation. Arai et al. target medium-range precipitation super-resolution over Japan (Arai et al. 2025), and Li et al. refine FuXi forecasts to 1 km over China (Li et al. 2025); both employ deterministic neural networks. Molinaro et al. (2026) show that diffusion-based probabilistic downscaling can transfer across heterogeneous upstream models up to 90 h for near-surface variables. Because precipitation is difficult to forecast even at 0.25◦ resolution, many learned refinement methods condition on the full atmospheric state rather than sharpening the model’s own coarse precipitation field (Rojas-Campos et al. 2023); CorrDiff, for instance, excludes coarse precipitation from its inputs (Mardani et al. 2025). By contrast, our setting is motivated by the fact that recent AIFS versions have improved precipitation skill enough for the coarse precipitation field itself to become a meaningful predictor (Moldovan et al. 2025). This does not imply that the AIFS precipitation forecast is unbiased; rather, it already contains useful information about event timing and broad spatial organization. At the same time, AIFS still exhibits lead-time-dependent amplitude errors, wet-area biases, and spatial displacements. Applying generative refinement directly to such a drifting forecast risks amplifying existing errors rather than correcting them, which motivates a debias-then-generate strategy (Wan et al. 2023). This motivates SwAIther-Precip, a two-step Swiss-AI precipitation downscaling framework that post-processes coarse-resolution precipitation forecasts from a global neural weather model (here AIFS, though the pipeline could generalize to other global weather model) into highresolution probabilistic precipitation fields over Switzerland. A key design choice is to introduce an explicit, lead-time-aware correction stage at coarse resolution before stochastic super-resolution. A unified deterministic correction network uses the full AIFS context and the forecast lead time to produce a physically realistic, intermediate, coarse precipitation field, and a diffusion-based model then generatively super-resolves that corrected precipitation field only. In particular, we do not attempt to autoregressively integrate a km-scale atmospheric state forward in time given the medium-range lead times of interest. Our main contributions are threefold: 1. A two-step downscaling decomposition. We present a pipeline that explicitly separates lead-time-aware
3 coarse-resolution bias correction from stochastic super-resolution, allowing each component to specialize: the deterministic first step corrects systematic spatial and intensity biases in the global model output, while the generative second step adds realistic fine-scale variability based on the observational training data. A key practical advantage is that the bias correction absorbs all lead-time dependence, so the super-resolution model operates on only two channels (corrected precipitation and topography) and can be trained in a perfect-prognosis setting on observations alone, substantially reducing computational cost. A proxy ablation suggests that single-step direct downscaling is substantially harder, though a complete comparison remains future work. We further identify a mesoscale variance bottleneck (20–100 km) introduced by the deterministic bias correction, pointing to a clear avenue for improvement. 2. Lead-time-aware downscaling. We introduce a single unified bias-correction model conditioned on forecast lead time through Feature-wise Linear Modulation (FiLM, Perez et al. (2018)) layers, covering lead times from 6 h to 6 days. This “multi-horizon” approach not only eliminates the need for separate lead-time-specific models, but also outperforms leadtime-specific specialists at longer lead times, suggesting that joint training across lead times acts as an effective form of regularization. 3. Km-scale probabilistic fields with realistic spatial structure. The full pipeline generates ensemble forecasts that reproduce the observed precipitation power spectrum down to an effective resolution of ∼4 km on a 1 km grid, with high spectral fidelity across all lead times up to 5 days. This is achieved while maintaining competitive probabilistic scores and well-calibrated ensemble spread. The remainder of the article is organized as follows. Section 2 describes the proposed methodology. Section 3 presents the datasets used for the study and for all experiments performed to test the methodology. Section 4 defines the baseline models, evaluation metrics, and analysis tools used to assess the methodology across multiple variants. Section 5 presents and analyzes results of the conducted experiments, at both the bias-correction and full downscaling levels. Section 6 summarizes the main findings and outlines future challenges and directions. 2. Methodology a. Problem formulation: lead-time-aware downscaling Let 𝑡 ∈ T denote the forecast valid time, i.e., the physical time at which the forecast is evaluated, and let ℓ ∈ L denote the lead time, i.e., the elapsed time since forecast
initialization. For each pair (𝑡, ℓ), let x𝑡 ,ℓ ∈ R 𝐻 ×𝑊 ×𝐶in denote the coarse-resolution multichannel forecast on a grid of height 𝐻 and width 𝑊, where the 𝐶in channels include precipitation, dynamic atmospheric variables (e.g., wind and humidity), and static fields (e.g., topography). Let y𝑡 ,ℓ ∈ R 𝑃𝐻 × 𝑃𝑊 ×𝐶out denote the corresponding highresolution reference field on a grid of height 𝑃 𝐻 and width 𝑃𝑊 , with 𝑃 𝐻 ≈ 𝑠𝐻 and 𝑃𝑊 ≈ 𝑠𝑊 for spatial upscale factor 𝑠. In this study, precipitation is the only target variable, so 𝐶out = 1, although the formulation naturally extends to multivariate outputs. The objective of lead-time-aware downscaling is to learn a parametrized mapping 𝑓 𝜃 : R 𝐻 ×𝑊 ×𝐶in × L → R 𝑃𝐻 × 𝑃𝑊 ×𝐶out , (1) that transforms coarse-resolution forecasts into highresolution precipitation fields while explicitly accounting for the forecast lead time. Unlike classical super-resolution, which aims solely at increasing spatial resolution, this mapping must also correct the systematic biases present in global model forecasts. Conditioning on ℓ is essential because forecast uncertainty, spatial displacement errors, and temporal decorrelation all increase with lead time, making the input–output relationship non-stationary across lead times. The potential use of multiple specialized singlelead-time models trained separately to achieve the same goal, instead of one lead-time conditioned model, will be discussed and compared throughout the study.
b y𝑡 ,ℓ = 𝑓 𝜃 (x𝑡 ,ℓ , ℓ),
b. Two-step lead-time-aware downscaling strategy We propose to decompose the lead-time-aware downscaling problem into two complementary sub-tasks: (i) Step 1: deterministic bias correction at coarse resolution, which maps the multi-channel global forecast x𝑡 ,ℓ ∈ R 𝐻 ×𝑊 ×𝐶in to an intermediate bias-corrected precipitation field e x𝑡 ,ℓ ∈ R 𝐻 ×𝑊 ×𝐶out , and (ii) Step 2: generative super-resolution, which refines e x𝑡 ,ℓ into high-resolution precipitation fields. This decomposition is motivated by the magnitude and structure of the systematic errors in global model forecasts, particularly at longer lead times, and allows each component to specialize: the bias correction handles large-scale amplitude and distributional mismatches, while the super-resolution focuses on fine-scale spatial variability. Three training strategies can implement this decomposition. In Strategy 1 (adopted in this work), the two steps are trained independently and applied sequentially. The bias correction model 𝑏 𝜙 produces: e x𝑡 ,ℓ = 𝑏 𝜙 (x𝑡 ,ℓ , ℓ),
(2)
and the super-resolution model 𝑠 𝜓 is trained in a perfect super-resolution setting using coarsened observations y𝑡↓,ℓ
. ..
2h ,1
6h
ad Le
ℓ
=
, tim . . . ,1 es 4
4h
··
·
Global model predictions
4
Dir
ect
Step 1: Lead-time aware Bias Correction
Ta sk
β (shift)
MLP
Low-res observation
et,ℓ = bϕ (xt,ℓ , ℓ) x
γ (scale)
egy
2)
ℓ
Step 2: Super-Resolution (CorrDiff) reg bt,ℓ y = rψ (e xt,ℓ )
Regression Net
mean cond.
Diffusion Net
+
reg diff bt,ℓ ∆yt,ℓ ∼ dω (e xt,ℓ , y )
reg diff bt,ℓ = y bt,ℓ y + ∆yt,ℓ
(E ensemble members)
High-res observation
Bias Correction Net
(St
rat
Lead time conditioning (FiLM)
Fig. 1: The SwAIther-Precip strategy: A first bias-correction step is followed by a super-resolution step (decomposed into regression and residual diffusion sub-step). During training, Step 1 and Step 2 are trained independently; during inference, a new global prediction (with a provided lead time) is first bias-corrected, and then passed through the super-resolution model to generate diffusion samples at high resolution.
as input:
c. Step 1: Lead-time-aware bias correction b y𝑡 ,ℓ ∼ 𝑠 𝜓 (y𝑡↓,ℓ ).
(3)
At inference, 𝑠 𝜓 is applied to the bias-corrected field e x𝑡 ,ℓ instead. Strategy 2 (used as ablation baseline) trains a single generative model to jointly perform bias correction and super-resolution, mapping x𝑡 ,ℓ directly to high-resolution output. Strategy 3 (left for future work) would train both steps end-to-end with joint loss terms, potentially mitigating error propagation from the sequential approach. We adopt Strategy 1. The reasons for this choice are threefold. First, decoupling the two steps allows us to explicitly isolate and evaluate the bias correction component, which constitutes the main methodological contribution of this work. Second, preliminary experiments show that leadtime-dependent biases in global forecasts hinder the training of a single end-to-end model (Strategy 2), whereas an explicit correction provides a cleaner input distribution for the diffusion model. Third, Strategy 1 offers computational advantages: the diffusion stage operates on biascorrected precipitation and static fields only, avoiding the memory overhead of processing all atmospheric channels. For completeness, we compare Strategy 1 to Strategy 2 at a representative 3-day lead time in Section 4.
The first step of SwAIther-Precip is a deterministic bias correction model operating at coarse resolution. It maps low-resolution multi-channel global model predictions to low-resolution observed precipitation fields, correcting systematic errors while preserving large-scale spatial structure. We maintain the same notations previously defined for the different quantities in the framework. 1) Bias correction model architecture The model is implemented as a convolutional U-Net with attention at lower levels, following the regression architecture of CorrDiff (Mardani et al. 2025). This choice aims to capture both large-scale spatial patterns and localized corrections while modeling long-range dependencies via attention, which are particularly relevant for precipitation over complex terrain. Lead-time conditioning via FiLM. To account for the strong dependence of forecast biases on lead time, we condition the model on ℓ via Feature-wise Linear Modulation (FiLM; Perez et al. 2018). The discrete lead time index is mapped to a continuous embedding eℓ , from which a small
5 MLP 𝜑 produces channel-wise scale and shift parameters ( γℓ , βℓ ) = 𝜑(eℓ ) applied to the input forecast as x̄𝑡 ,ℓ = FiLM(x𝑡 ,ℓ | ℓ) = x𝑡 ,ℓ ⊙ (1 + γℓ ) + βℓ ,
(4)
where ⊙ denotes channel-wise multiplication. In practice, FiLM conditioning proved more effective than appending lead time as an additional input channel. Residual formulation. To anchor predictions to the original forecast and encourage the network to learn systematic corrections rather than reconstruct the full field, we adopt a (prec) residual formulation. Denoting by 𝑝¯0 = x̄𝑡 ,ℓ the precipitation channel of the FiLM-modulated forecast, the final bias-corrected output is: e x𝑡 ,ℓ = 𝑝¯0 + 𝛾res · UNet 𝜙 x̄𝑡 ,ℓ (5) where 𝛾res is a learned scalar initialized at zero. Anchoring (prec) on 𝑝¯0 rather than on the raw forecast precipitation x𝑡 ,ℓ provides a useful inductive bias: when the U-Net residual is uninformative, the model reduces to a lead-aware affine recalibration of the raw forecast rather than to an identity copy, generalizing the standard residual setup (He et al. 2016) to a forecast-horizon-aware baseline. Output values are clipped to [0, 20] mm/6 h to enforce non-negativity and provide a numerical safeguard during training (cap value discussed in the Supplementary Material, Section S3). The detailed architecture of the U-Net and the full Step 1 pipeline are presented in Appendix b. 2) Regularization for cross-lead-time consistency Multiple forecasts with different lead times may correspond to the same valid time 𝑡, in which case the biascorrected field should be approximately invariant to ℓ for this given 𝑡. We enforce this via a regularization term penalizing discrepancies across lead times: Ltime =
∑︁ 1 2 e x𝑡 ,ℓ −e x𝑡 ,ℓ ′ 2 , 2 |L𝑡 | ℓ,ℓ ′ ∈ S 𝑡∈T ∑︁
(6)
𝑡
where S𝑡 is the set of available lead times for a valid time 𝑡. This regularization term encourages consistency across lead times while still allowing the model to adapt to systematic differences in forecast skill as a function of ℓ. It acts as a form of temporal smoothing and reduces spurious lead-time-dependent artifacts in the bias-corrected fields. 3) Bias-correction training objective The bias correction model is trained using a combination of a reconstruction loss at coarse resolution and the temporal regularization term introduced above. For each training sample (𝑡, ℓ), we define a Huber loss between the biascorrected output and the corresponding low-resolution observed precipitation:
2 (𝑖) ↓, (𝑖) ↓, (𝑖) 1 ∑︁ e 𝑥 − 𝑦 , if e 𝑥 𝑡(𝑖) ≤ 𝛿, 1 2 𝑡 ,ℓ 𝑡 ,ℓ ,ℓ − 𝑦 𝑡 ,ℓ LH (𝑡, ℓ) = ↓, (𝑖) 𝑁 𝑖 𝑥 𝑡(𝑖) − 21 𝛿 ,otherwise, 𝛿 e ,ℓ − 𝑦 𝑡 ,ℓ (7) where 𝛿 is the Huber threshold, 𝑁 = 𝐻 × 𝑊 × 𝐶out is the total number of elements, and y𝑡↓,ℓ ∈ R 𝐻 ×𝑊 ×𝐶out denotes the observed precipitation field coarsened to the AIFS grid. Unless specified, we choose 𝛿 = 1. The Huber loss acts as mean squared error for small residuals and transitions to a mean absolute error for large residuals, making the training more robust to outliers in the precipitation field. The total training loss for the bias correction U-Net is then given by Lbias = E (𝑡 ,ℓ )∼D [LH (𝑡, ℓ)] + 𝜆 Ltime ,
(8)
where 𝜆 is a hyperparameter controlling the strength of the cross-lead-time regularization. 4) Curriculum learning and multi-step fine-tuning Training a single model across all lead times is challenging because error characteristics vary strongly between short-range and medium-range forecasts. To improve convergence across lead times, we explore two progressive training strategies , both based on the idea of gradually exposing the model to longer (and more difficult) lead times. Curriculum learning Bengio et al. (2009) operates within a single training phase: batches are filtered to include only a subset of lead times that expands at fixed epoch intervals (e.g., lead times 6–48 h for the first 4 epochs, then 48–96 h, etc.). Multi-step fine-tuning Siddiqui et al. (2024) organizes training into multiple stages with separate hyperparameters. We adopted a three stages scheme. Stage 1 trains on short-range lead times only (first third of lead times) with 𝜆 = 0, establishing a strong spatial representation. Stage 2 fine-tunes from the best Stage 1 checkpoint on mediumto-long-range lead times, still with 𝜆 = 0, to adapt to the distinct bias patterns at longer leads. Stage 3 fine-tunes on all lead times jointly with a smaller learning rate and 𝜆 > 0, harmonizing behavior across horizons while preserving the lead-specific corrections learned in earlier stages. d. Step 2: Generative super-resolution The second step of SwAIther-Precip focuses on spatial super-resolution using the CorrDiff framework (Mardani et al. 2025), which combines a deterministic regression model with a conditional diffusion model. Starting from a low-resolution precipitation field that has already been bias-corrected, the goal is to reconstruct realistic highresolution precipitation patterns consistent with observed fine-scale structure. Crucially, this stage is not conditioned
6 on forecast lead time: all lead-time-dependent effects are handled by the upstream bias correction, so the superresolution model learns to refine spatial structure independently of the forecast horizon.
e. Summary of the end-to-end inference pipeline Figure 1 summarizes the full inference pipeline. Given a coarse-resolution forecast x𝑡 ,ℓ , the pipeline applies the following sequential transformations:
1) Super-resolution model architecture
e x𝑡 ,ℓ = 𝑏 𝜙 (x𝑡 ,ℓ , ℓ),
The architecture of the super-resolution layer is the same as defined by Corrdiff (Mardani et al. 2025). The regression net 𝑟 𝜓 is a UNet which upsamples the low-resolution input (precipitation and topographic elevation only) to proreg duce an initial high-resolution estimate b y𝑡 ,ℓ = 𝑟 𝜓 (y𝑡↓,ℓ ). A conditional diffusion model 𝑑 𝜔 then generates stochastic residual corrections to restore fine-scale variability that the deterministic regression cannot capture:
b y𝑡 ,ℓ = 𝑟 𝜓 (e x𝑡 ,ℓ ) + Δydiff 𝑡 ,ℓ ,
reg b y𝑡 ,ℓ = b y𝑡 ,ℓ + Δydiff 𝑡 ,ℓ ,
reg
Δydiff x𝑡 ,ℓ ,b y𝑡 ,ℓ ). 𝑡 ,ℓ ∼ 𝑑 𝜔 (e
(10) reg Δydiff x𝑡 ,ℓ ,b y𝑡 ,ℓ ). 𝑡 ,ℓ ∼ 𝑑 𝜔 (e
(11)
where 𝑏 𝜙 is the lead-time-conditioned bias correction UNet, 𝑟 𝜓 the regression U-Net, and 𝑑 𝜔 the conditional diffusion model. This decomposition explicitly separates leadtime-dependent bias correction from lead-time-agnostic spatial refinement. 3. Data
(9)
Both models are trained independently following the standard CorrDiff protocol. 2) Training and inference configurations During training, the super-resolution model operates in a perfect super-resolution setting: low-resolution inputs are obtained by coarsening high-resolution observations, isolating the spatial refinement task from upstream forecast biases. At inference time, the model is instead applied to the bias-corrected fields e x𝑡 ,ℓ produced by Step 1. This separation ensures that the super-resolution model never compensates for global model biases. 3) Generative inference pipeline Since the super-resolution model is not lead-time-aware, inference is performed independently per lead time: for each lead time, the regression and diffusion networks are applied sequentially to the bias-corrected inputs. For each sample, 𝐸 ensemble members are generated using independent random seeds, producing stochastic realizations that share large-scale structure but differ in fine-scale detail. We note that sharing diffusion seeds across lead times has been proposed to encourage temporal coherence between forecast horizons (Andrae et al. 2024); however, in our setting, independent seeds yielded better results, likely because the bias correction step already provides leadtime-specific conditioning that makes artificial seed-based coherence unnecessary. The disadvantage of our present framework is the lack of temporal consistency between individual ensemble members: the super-resolution samples are independent when conditioned on the bias-correction outputs. Model outputs, generated in log-transformed and normalized space, are transformed back to physical units (mm/6h) and clipped to non-negative values.
a. Target precipitation data The target dataset for the downscaling task is the MeteoSwiss CombiPrecip product (Sideris et al. 2014). This dataset combines measurements from a national weather radar system and rain gauge observations distributed across Switzerland. The radar–gauge combination provides the best estimate of ground-level sub-daily precipitation distribution currently available for Switzerland (as claimed by MeteoSwiss, https://www.meteoswiss.admin.ch/). The original CombiPrecip data are expressed in units of mm/h. The dataset provides precipitation fields at an approximate spatial resolution of 1 km and a temporal resolution of 10 minutes, and is available continuously from 2005 onward. The dataset considers a spatial windows covering Switzerland and its immediate surroundings, corresponding to a grid of 640 × 710 pixels. To ensure temporal consistency with the global neural weather model forecasts, all CombiPrecip fields are aggregated into 6-hour accumulated precipitation amounts (mm/6h) by summing the corresponding high-frequency observations. This temporal aggregation aligns the observational targets with the forecast horizons considered in this work. For the bias correction task, the same CombiPrecip data are spatially aggregated to match the coarse resolution of the global neural weather model forecasts. This coarsened version of the observations serves as the low-resolution target for the deterministic bias correction model. The aggregation is performed such that the resulting grid corresponds approximately to a 31 km resolution, consistent with typical global forecasting systems. b. Input forecast data 1) AIFS global forecast inputs for bias correction The input data for the bias correction stage are forecasts from ECMWF’s Artificial Intelligence Forecast-
7
AIFS predictions ( 31 km)
(a)
Temperature (500 hPa)
(b)
53900 54000 54100 54200 54300 54400 54500 54600 m²/s²
248.0 248.5 249.0 249.5 250.0 250.5 251.0 251.5 K
(e)
(f)
0.0
CombiPrecip ( 1 km)
Geopotential (500 hPa)
Total precipitation
0.2
1.0
3.0 mm/6h
7.0
30.0
0.0
0.2
1.0
3.0 mm/6h
7.0
30.0
Total Precipitation
(i)
0.0
Convective precipitation
0.2
1.0
3.0 mm/6h
7.0
(c)
Specific humidity (500 hPa)
0.0004 0.0005 0.0006 0.0007 0.0008 0.0009 0.0010 kg/kg
10 m wind speed
(g)
0
1
0
m/s
3
0.0
(h)
4
0
Total cloud cover
0.2
0.4
0.6
0.8
500
500 1000 1500 2000 2500 3000 3500 4000 m
1000 1500 2000 2500 3000 3500 4000 m
Fig. 2: Overview of the input and target data for a choice of time step. (a–d) AIFS forecast fields at ∼31 km resolution used as input to the bias correction model: geopotential, temperature, and specific humidity at 500 hPa, and total cloud cover. (e–h) Additional AIFS surface-level inputs: total and convective precipitation, 10 m wind speed, and model orography. Note that the geopotential, temperature and specific humidity at 850 hPa were also added to the list of used AIFS outputs. (i, j) CombiPrecip high-resolution observational target at ∼1 km: 6 h accumulated precipitation and terrain altitude. The resolution gap between the coarse AIFS grid (23×37 points) and the CombiPrecip target (640×710 points) corresponds to a ∼30× downscaling factor, illustrating the challenge of recovering fine-scale precipitation structure from the global model output.
ing System (AIFS; Lang et al. 2024), specifically the AIFS-single-1.0 configuration, which offers improved precipitation skill over earlier versions through physicsinformed constraints (Moldovan et al. 2025). The available ensemble configuration was not used, as its deterministic precipitation scores are lower than those of the single-member model. Forecasts were generated using the official implementation1 with IFS analysis initial conditions, covering the period 2019–2023 at lead times ℓ ∈ {6h, 12h, . . . , 144h}. From the large set of available AIFS outputs (the full list can be found in the Supplementary Material in Section S1), we retain a subset of physically relevant variables summarized in Figure 2 (with geopotential, temperature and specific humidity at 850 hPa additionally included).
1 https://huggingface.co/ecmwf/aifs-single-1.0
1.0
Altitude
Altitude
(j)
30.0
2
(d)
This subset captures the key physical drivers of precipitation over complex terrain: precipitation fields provide the direct correction signal, near-surface winds inform orographic forcing, cloud cover indicates the condensation regime, and upper-air variables at 500 and 850 hPa characterize large-scale steering flow, atmospheric stability, and moisture availability - two levels established as among the most informative for statistical precipitation downscaling (Hessami et al. 2008; Wilby et al. 2002). The 850 hPa level is particularly relevant over Switzerland, where it intersects the Alpine crest and captures low-level moisture transport. Retaining a compact input set also keeps dimensionality manageable and limits overfitting to redundant channels. All input fields are interpolated to a common coarse spatial grid covering Switzerland, with a spatial resolution of approximately 31 km. Forecast fields are temporally aligned
8 with the corresponding CombiPrecip observations at each valid time and lead time.
and (iii) held-out time windows reflect intrinsic difficulty rather than a generalization gap.
2) Inputs to the super-resolution model
d. Data preprocessing and postprocessing
Both the regression and diffusion stages of the super resolution step operate on a restricted set of input channels. Specifically, the low-resolution inputs consist only of precipitation and static topographic information. We experimented with including all additional atmospheric variables predicted by the global neural model and used in the bias correction model, but found no consistent performance improvements at the spatial scales considered here. Consequently, these variables are omitted from the superresolution stage. The output of the super-resolution model is the high-resolution precipitation field.
1) Preprocessing of AIFS precipitation forecasts The precipitation outputs produced by the global neural weather model (AIFS in our case) are first converted to a common physical unit and temporal resolution. All precipitation forecasts are expressed as accumulated precipitation in millimeters over a 6-hour interval (mm/6h), consistent with the evaluation horizons considered in this study. Since negative precipitation values are physically meaningless and occasionally arise due to numerical artifacts in neural model outputs, all negative values are clipped to zero. This operation is applied consistently across all lead times and forecast instances.
c. Weekly Train-Test split We adopt a rolling weekly Train-Test split strategy inspired by Google’s MetNet training scheme (Sønderby et al. 2020; Espeholt et al. 2022), which employed a similar interleaved temporal splitting approach for precipitation forecasting. For each calendar week, the first six days are assigned to training and the seventh to testing; all lead times sharing a given valid time belong to the same split. Within each day, only the 06–12 UTC and 18–00 UTC accumulation windows are used, which introduces a minimum six-hour gap to the held-out windows (00–06 UTC and 12–18 UTC), reducing residual temporal autocorrelation while preserving the unused windows as an independent testbed for temporal generalization. We do not reserve a separate validation set for hyperparameter tuning. Given the limited temporal coverage of the dataset (2019–2023), we prioritize maximizing both training and test data volume. Hyperparameters (learning rate, number of epochs, 𝜆, curriculum schedule) were set based on standard practices in the literature and brief preliminary experiments; no systematic optimization (e.g., grid search or Bayesian tuning) was performed against test set metrics. The comparison of multiple training configurations (Section 5) should therefore be interpreted as an architectural and methodological study rather than a hyperparameter search. This weekly splitting strategy preserves a realistic seasonal distribution of weather regimes in both splits and avoids the sensitivity to interannual distribution shifts that can affect year-based partitions. Additional experiments validating these design choices - including autocorrelation analysis, evaluation on held-out time windows, and comparison with a yearly split - are presented in the Supplementary Material (Section S4), confirming that (i) temporal autocorrelation is not a significant confound, (ii) inter-year performance differences reflect distribution shift rather than leakage,
2) Preprocessing of CombiPrecip precipitation observations The MeteoSwiss Combi precipitation product is available at a higher temporal resolution than the global model forecasts. To ensure temporal alignment, the observations are aggregated into 6-hour accumulated precipitation fields (mm/6h) by summing the corresponding hourly values. To avoid training the model on uninformative samples, we remove dry samples, defined as precipitation fields in which more than 99.5% of grid points exhibit zero precipitation. Such samples provide little learning signal for either bias correction or super-resolution and disproportionately dominate the dataset due to the sparsity of precipitation events. In addition, we mitigate the impact of extreme outliers, which are likely attributable to measurement or processing errors. For each precipitation field, we fit a Gamma distribution, a commonly used parametric model for precipitation, to the non-zero pixel values. Let 𝐹Γ denote the cumulative distribution function of the fitted Gamma distribution. We compute the 99.5th percentile threshold 𝜏 = 𝐹Γ−1 (0.995), and clip all precipitation values exceeding this threshold. This procedure preserves the overall structure of intense precipitation events while preventing a small number of extreme values from dominating the learning process. 3) Shared normalization and transformation Following the above steps, both input and target precipitation fields are transformed using a logarithmic variancestabilizing transformation: 𝑥 ← log(1 + 𝑥), which reduces skewness and mitigates the influence of large precipitation values.
9 Table 2: Summary of evaluation metrics
Table 1: Baseline models for bias correction evaluation Model
Description
Category
Metric
Range
Optimal
B0 (raw AIFS) B1–B3
Point-wise
MSE MSEwet
[0, ∞) [0, ∞)
0 0
Error magnitude Error for rain events
B4 B5 B7
Unprocessed AIFS precipitation prediction Affine corrections of increasing granularity: global (B1), per-lead (B2), per-pixel per-lead (B3) Per-pixel multivariate linear regression (ridge) Per-pixel per-lead quantile mapping Global multivariate ridge regression
Categorical
CSI
[0, 1]
1
Overall detection skill
Spatial
ViT + FiLM CorrDiff reg. Single-lead UNets
Vision Transformer with lead-time conditioning Regression-only direct downscaling (Strategy 2 proxy) Independent UNets at 6 h, 3 days and 6 days
FSS SAL-S SAL-A SAL-L
[0, 1] [ −2, 2] [ −2, 2] [0, 2]
1 0 0 0
Scale-dependent skill Structure error Amplitude error Location error
Ensemble
CRPS avFSS eS eA eL
[0, ∞) [0, 1] [ −2, 2] [ −2, 2] [0, 2]
0 1 1 1 1
Overall probabilistic skill Scale-dependent skill Median Structure error Median Amplitude error Median Location error
Finally, precipitation fields are normalized using min–max scaling computed over the training dataset, where 𝑥min and 𝑥max are determined exclusively from training samples and applied consistently during test and inference. This normalization ensures numerical stability during optimization and consistent scaling across datasets. 4) Postprocessing of downscaled forecasts Throughout this study, a wet/dry threshold of 0.1 mm/6h is used, as 0.1mm corresponds to the minimum reporting resolution of WMO-standard rain gauges (WMO, 2018, No. 8, Ch. 6), regardless of the accumulation time. The high-resolution diffusion samples undergo two postprocessing steps: (1) values are masked to regions where the bias-corrected low-resolution field indicates precipitation, enforcing spatial consistency between the two pipeline stages; (2) the 0.1 mm/6h threshold is applied to suppress residual drizzle artifacts. The same threshold is applied to observations for consistent evaluation. At the bias correction stage, no thresholding is applied, as the untouched field is passed directly as input to super-resolution. 4. Evaluation framework a. Baseline models To quantify the benefits of the proposed two-step approach, we compare against classical post-processing baselines and architectural alternatives, summarized in Table 1. All baselines are evaluated at coarse resolution for the bias correction task; for the full downscaling pipeline, only a selected subset is evaluated since performance is dominated (prec) by the upstream bias correction quality. Let x𝑡 ,ℓ ∈ R 𝐻 ×𝑊 denote the raw coarse-resolution precipitation forecast and y𝑡↓,ℓ ∈ R 𝐻 ×𝑊 the corresponding observational target. The deep learning baselines, which require more details, are as follows: Vision Transformer (ViT) with FiLM conditioning. An architectural alternative to the U-Net: the FiLM-conditioned input is split into non-overlapping patches and processed by a Transformer encoder with multi-head self-attention,
Purpose
and decoded via transposed convolutions to reconstruct the coarse-resolution precipitation field CorrDiff regression-only (Strategy 2 proxy). The regression component of CorrDiff trained directly from coarse forecasts to high-resolution precipitation without diffusion refinement, testing whether a single-step deterministic mapping can substitute for the two-step decomposition. Single-lead U-Net models. Separate U-Net models trained independently at 6 h, 3 days, and 6 days, sharing the same architecture but without lead-time conditioning, to assess the benefit of multi-horizon training. b. Forecast verification metrics We evaluate bias correction and downscaling performance using metrics spanning four complementary aspects of forecast quality. Full definitions of metrics and implementation details are provided in the Supplementary Material (Section S6). Pointwise and categorical metrics. We report MSE restricted to wet pixels (MSEwet , with a threshold 0.1 mm/6h, as mentioned in Section 3d) to focus on precipitation events rather than the dominant dry fraction, and the Critical Success Index (CSI) for binary event detection skill. Spatial verification. The Fractions Skill Score (FSS; Roberts and Lean 2008) evaluates spatial agreement at multiple neighborhood sizes (𝑛 = 1, 5, 17, corresponding to pixel-level, ∼25 km, and ∼85 km scales). For ensemble forecasts we use the member-averaged FSS (avFSS; Necker et al. 2024). The Structure–Amplitude–Location diagnostic (SAL; Wernli et al. 2008) decomposes forecast error into shape, intensity, and displacement components; for ensembles we report the median across members (eSAL; Radanovics et al. 2018). Probabilistic verification. The Continuous Ranked Probability Score (CRPS; Hersbach 2000) provides a strictly proper measure of ensemble skill. Ensemble calibration
10 is assessed via randomized Probability Integral Transform histograms (Gneiting and Raftery 2007), with departure from uniformity quantified by KL divergence.
a. Results for lead-time-aware bias correction (step 1)
Spectral evaluation. Following Sinclair and Pegram (2005); Pulkkinen et al. (2019), we compute radiallyaveraged power spectral densities of co-masked logprecipitation fields and report band-averaged spectral ratios (predicted/observed PSD) at large (100–600 km), meso (20–100 km), and small (2–20 km) scales. The effective resolution is defined as the wavelength where the spectral ratio drops below 0.5 (Klaver et al. 2020).
Figure 3 summarizes all metrics averaged across lead times. Without post-processing, raw AIFS (B0) yields average MSE wet of 0.088. Global affine corrections (B1– B2) do not improve MSE and severely distort precipitation structure (SAL-S ≈ −1.38, SAL-A ≈ −1.54). Among classical baselines, the strongest is B4 (per-pixel multivariate regression), achieving MSE wet of 0.073 - a 17% reduction over B0 - highlighting the importance of both local calibration and additional atmospheric predictors. All learned models substantially outperform B4: the best variant, UNet Multi FT, reaches average MSE wet of 0.065 (26% reduction over B4). Combining curriculum learning with temporal regularization (UNet𝜆=2 , Curri) does not improve over UNet𝜆=2 alone (0.072 vs. 0.066), suggesting partial redundancy between the two mechanisms. The ViT achieves 0.071, comparable to UNet𝜆=0 but behind the best UNet variants.
c. Verification protocol and Implementation details All metrics are computed separately for each lead time, allowing us to assess how model performance degrades with increasing forecast horizon. We evaluate 24 lead times ranging from 6 hours to 6 days (144 hours) at 6-hour intervals. Additionally, we report metrics averaged across all lead times to provide a summary measure of overall model performance. For all probabilistic metrics, we generate 𝑀 = 12 ensemble members from the diffusion model for each forecast initialization. Metrics are computed at each lead time separately and additionally averaged across all leads. Table 2 summarizes all evaluation metrics used in this study, their optimal values, and their primary purpose. 5. Results This section reports results for (i) the coarse-resolution bias correction task and (ii) the full downscaling pipeline. All scores follow the evaluation protocol described in Sections 3c and 3d, and are computed per lead time unless stated otherwise. We evaluate several training configurations of the leadtime-aware bias correction U-Net, varying the temporal regularization strength 𝜆 (Section c), the use of curriculum learning, and whether multi-step fine-tuning is applied (Section 4). Model names encode these choices: e.g., UNet𝜆=2 , Curri denotes 𝜆 = 2 with curriculum learning and no multi-step fine-tuning; UNet Multi FT denotes UNet𝜆=2 further fine-tuned with a multi-step loss. We also include a ViT baseline trained with the same protocol as UNet𝜆=0 , Curri. All multi-lead models are conditioned on lead time via FiLM and trained jointly across all 24 forecast horizons (6 h to 144 h). Single-horizon U-Nets trained independently at 6 h, 3 days, and 6 days, as well as a Strategy 2 baseline (Section 4), are included for comparison. The baselines include the non deep learning baselines mentioned in Section 4a: raw AIFS output (B0), global affine calibration with a single scaling (B1) or per-lead scaling (B2), per-pixel linear regression on precipitation only (B3), per-pixel multivariate linear regression (B4), quantile mapping (B5), and global ridge regression (B7).
1) Performance averaged across lead times
For event detection and spatial skill, the pattern is consistent. UNet𝜆=0 , Curri achieves the highest average CSI (0.541), while UNet𝜆=2 , Curri leads in FSS5 (0.832). All learned models exceed 0.81 in FSS5 , compared to 0.800 for B4 and 0.607 for raw AIFS. Interestingly, CSI rankings differ slightly from MSE: UNet𝜆=0 , Curri leads in CSI despite not being the best in MSE, suggesting that curriculum learning may benefit event detection more than squared error reduction. The SAL diagnostic (Figure 3) confirms these trends: raw AIFS shows strong positive structure and amplitude biases (overly broad and intense fields), which spatially varying baselines (B3–B4) and all learned models reduce to near-neutral values. UNet Multi FT stands out with a near-zero structure and the only positive amplitude bias, foreshadowing the bias-variance trade-off analyzed in Section 6. Location errors remain small across all learned models (SAL-L ≤ 0.027). 2) Lead-time dependence of skill Figure 5 provides a comprehensive view of bias correction performance across all lead times, clearly illustrating the consistent advantage of learned models over classical baselines at every horizon. Figure 4 further details MSE wet, CSI, and FSS5 at representative lead times. All models degrade with increasing horizon, but Unet models maintain a consistent advantage over baselines across all leads (Figures 5). UNet𝜆=2 achieves the lowest MSE at short leads (0.048 at 6 h), while UNet Multi FT is strongest at extended ranges (0.085 at 6 d vs. 0.090 for UNet𝜆=2 ), suggesting that multi-step fine-tuning improves long-range error accumulation. The ViT is competitive at the longest horizons (CSI of 0.432 at 6 d), suggesting robustness to increased input uncertainty at extended ranges.
11 B0: AIFS (raw)
0.050
0.088
0.355
0.607
+0.830
+1.188
0.062
B1: Global Affine
0.024
0.096
0.412
0.695
-1.383
-1.545
0.085
B2: Affine / Lead
0.024
0.096
0.411
0.695
-1.379
-1.534
0.090
B3: Pixel LR (precip)
0.020
0.084
0.470
0.766
-0.176
-0.117
0.019
B4: Pixel LR (multivar)
0.018
0.073
0.508
0.800
-0.169
-0.108
0.015
B5: Quantile Mapping
0.040
0.135
0.347
0.792
-0.733
-0.001
0.025
B7: Ridge Reg Global
0.023
0.095
0.430
0.722
-0.873
-0.899
0.029
UNet =0, Curri
0.016
0.068
0.541
0.830
-0.230
-0.275
0.022
UNet =0
0.018
0.074
0.510
0.811
-0.301
-0.350
0.027
UNet =2
0.016
0.066
0.539
0.828
-0.188
-0.197
0.020
UNet Multi FT
0.017
0.065
0.535
0.821
+0.017
+0.195
0.021
UNet =2, Curri
0.016
0.072
0.540
0.832
-0.334
-0.394
0.023
ViT
0.017
0.071
0.538
0.822
-0.006
+0.031
0.026
MSE (all) Better
MSE (wet)
CSI
FSS
SAL-S
SAL-A
SAL-L
Worse
(colors normalized independently per column; for SAL-S and SAL-A, closer to 0 is better)
Fig. 3: Bias Correction (Step 1): General model comparison across all leads, for various metrics of interest (averaged across the 24 considered 6-hour leads) Colors are normalized independently per column. For SAL-S and SAL-A, values closer to zero indicate better performance.
3) Multi-lead training versus single-lead specialization To assess whether explicit lead conditioning is preferable to training separate horizon-specific models, we compare lead-aware models to single-lead UNets trained independently at 6 h, 3 d, and 6 d (Figure 4). The advantage of multi-lead training grows with lead time. At 6 h, the single-lead specialist achieves slightly higher CSI (0.636 vs. 0.613) and FSS5 (0.904 vs. 0.898) while matching the best multi-lead model in MSE (0.049). At 3 days, performance is comparable across approaches. At 6 days, however, the multi-lead models are clearly superior: the single-lead UNet reaches only 0.100 MSE and 0.408 CSI, compared to 0.085 and 0.429 for UNet Multi FT. This pattern suggests that specialization provides marginal benefits at short horizons, but that the broader training distribution of multi-lead models provides regularization that is increasingly beneficial at longer lead times. Operationally, a single unified model that remains competitive with or surpasses dedicated specialists is also far more practical than maintaining a separate model per lead time.
4) Two-step decomposition versus approximate direct downcaling To probe whether the explicit bias correction step is beneficial, we compare against a CorrDiff regression model trained to map AIFS fields directly to high-resolution precipitation (Strategy 2), bypassing the intermediate correction. This comparison is approximate: we use only the regression component of CorrDiff as a proxy, since running the full diffusion on all atmospheric input channels would drastically increase memory requirements and require spatial patching that introduces boundary artifacts, making a fair comparison difficult. Nevertheless, the regression component captures the deterministic mapping capacity of Strategy 2 and provides a useful lower bound on its performance. We further note that conditioning the Step 2 regression on all atmospheric channels (rather than just precipitation and altitude) yielded negligible improvement within Strategy 1, suggesting that coarse-resolution atmospheric fields add little information at the 1km target scale. The results favor the two-step approach. At 72 h, Strategy 2 achieves MSE of 0.093 (vs. 0.062–0.065 for the best bias correction models), CSI of 0.344 (vs. 0.557–0.560), and FSS5 of 0.647 (vs. 0.842–0.846). The SAL diagnostic reveals both structural distortion (SAL-S = −0.421)
12
Baselines Multi-lead (ours)
B3: Pixel LR (precip)
B2: Affine / Lead
B1: Global Affine
B0: AIFS (raw)
0.115
0.059
0.070
0.082
0.087
0.089
0.086
0.116
0.058
0.069
0.081
0.087
0.086
0.051
0.086
0.117
0.058
0.070
0.081
0.086
0.086
0.050
0.085
0.118
0.058
0.069
0.081
0.086
0.086
0.054
0.088
0.122
0.061
0.073
0.085
0.089
0.083
0.062
0.093
0.133
0.068
0.080
0.093
0.094
0.083
0.071
0.098
0.142
0.076
0.087
0.099
0.099
0.089
0.083
0.103
0.151
0.086
0.096
0.109
0.104
0.094
0.097
0.113
0.165
0.100
0.110
0.123
0.115
0.101
0.068
0.095
0.135
0.073
0.084
0.096
0.096
0.088
0.613 0.619 0.610 0.603 0.587 0.557 0.520 0.488 0.432 0.541
0.476 0.475 0.474 0.472 0.460 0.438 0.417 0.392 0.354 0.430
0.400 0.416 0.410 0.395 0.387 0.365 0.331 0.290 0.243 0.347
0.561 0.561 0.559 0.558 0.544 0.519 0.493 0.464 0.423 0.508
0.517 0.516 0.515 0.514 0.500 0.481 0.457 0.434 0.396 0.470
0.450 0.452 0.450 0.451 0.439 0.420 0.400 0.378 0.345 0.411
0.454 0.456 0.453 0.454 0.440 0.420 0.399 0.376 0.343 0.412
0.395 0.392 0.389 0.384 0.369 0.358 0.344 0.332 0.312 0.355
0.895 0.900 0.892 0.886 0.874 0.842 0.809 0.785 0.726 0.830
0.781 0.778 0.773 0.769 0.758 0.731 0.707 0.675 0.631 0.722
0.848 0.851 0.853 0.847 0.834 0.810 0.780 0.738 0.679 0.792
0.857 0.854 0.852 0.849 0.836 0.811 0.784 0.754 0.712 0.800
0.820 0.816 0.815 0.814 0.799 0.777 0.751 0.728 0.687 0.766
0.746 0.744 0.741 0.741 0.728 0.705 0.679 0.658 0.619 0.695
0.751 0.749 0.745 0.744 0.730 0.705 0.678 0.653 0.613 0.695
0.670 0.661 0.655 0.647 0.626 0.609 0.590 0.575 0.550 0.607
FSS5
B4: Pixel LR (multivar)
0.087 0.048
CSI
B5: Quantile Mapping
0.049
MSEwet
UNet =0, Curri 0.071
0.872 0.872 0.866 0.858 0.848 0.822 0.787 0.749 0.714 0.811
0.059
0.576 0.568 0.571 0.559 0.550 0.525 0.490 0.449 0.416 0.510
0.060
0.074
0.058
0.095
0.065 0.069
0.078
0.055 0.065
0.695
6d
Avg
0.074
UNet =0 0.051
0.099
5d
0.893 0.899 0.891 0.889 0.867 0.844 0.813 0.769 0.716 0.828
0.050 0.073
0.095
0.853
3d
0.647
4d
0.609 0.615 0.608 0.600 0.581 0.556 0.523 0.480 0.427 0.539
0.053 0.068
0.085
2d
0.066
0.050 0.065
0.079
1d
0.090
0.048 0.053 0.072
12h 18h
0.074
UNet =2 0.050
0.073
6h
0.888 0.888 0.870 0.879 0.867 0.844 0.801 0.769 0.718 0.821
0.051 0.069
B7: Ridge Reg Global
Worse
0.610 0.609 0.587 0.597 0.583 0.560 0.514 0.479 0.429 0.535
0.049
0.071
0.408
6d
Avg
0.065
0.050 0.058
5d
0.085
UNet Multi FT
0.059
0.568
0.344
3d
4d
0.898 0.903 0.894 0.890 0.872 0.846 0.803 0.783 0.728 0.832
0.057
2d
0.884 0.884 0.884 0.878 0.867 0.838 0.800 0.776 0.718 0.822
0.057
1d
0.613 0.618 0.608 0.601 0.583 0.559 0.514 0.483 0.431 0.540
0.058
12h 18h
0.608 0.606 0.602 0.593 0.585 0.557 0.515 0.492 0.432 0.538
0.059
6h
0.072
0.055
0.100
6d
(colors normalized independently per column)
Avg
0.071
0.056
5d
0.904
4d
0.636
0.056
3d
0.093
0.066
0.056
2d
UNet =2, Curri
1d
0.049
18h
ViT
12h
UNet Single 6h UNet Single 3d UNet Single 6d UNet Strat2 3d
6h
Better
Fig. 4: Bias correction (Step 1): MSE (wet pixels), CSI, and FSS5 by lead time. Models are grouped into classical baselines (top), multi-lead learned models (middle), and single-horizon / ablation models (bottom). Colors are normalized independently per column to highlight relative model ranking within each lead time and metric.
Single / Ablation
13
Fig. 5: Bias correction (Step 1): MSE over wet pixels (> 0.1 mm/6h) as a function of forecast lead time. Results are shown for multi-lead UNet and ViT variants, all baselines, and single-lead specialist models evaluated at their respective target leads.
and amplitude over-prediction (SAL-A = 0.594). While a definitive comparison would require a full end-to-end Strategy 2 pipeline, these results suggest that simultaneously correcting biases, bridging the resolution gap, and learning convective-scale structure is considerably harder than decomposing the problem into specialized steps. Beyond performance, the two-step decomposition offers a clear practical advantage: the super-resolution stage requires only two input channels and trains on observations directly, making it substantially cheaper to train and run than a single model operating on the full atmospheric state. 5) Qualitative examples of bias-corrected precipitation Figure 6 presents bias-corrected fields from four multilead models at 18 h and 72 h lead times, for examples of time steps in the test dataset. At 18 h, all models successfully correct the spatial over-extension of AIFS precipitation, recovering tighter spatial structures consistent with the ground truth. UNet𝜆=2 and Multi FT reproduce the sharpest spatial gradients for localized events, consistent with their leading MSE wet scores. At 72 h, predictions are visibly smoother but all models substantially reduce the large AIFS biases. Multi FT maintains relatively sharp features at this longer horizon, consistent with its nearzero SAL-A and best MSE wet at 3 days. Across both
lead times, differences between neural models remain subtle compared to the large shared improvement over the AIFS and classical baselines, consistent with the tight metric spread among learned models (CSI: 0.510–0.541; Figure 3). b. Results for the full downscaling pipeline (steps 1+2) We now evaluate the full two-step downscaling pipeline, where the bias-corrected low-resolution precipitation is refined using CorrDiff-based generative super-resolution. Due to the computational cost of diffusion sampling - generating 12 ensemble members per time step across all lead times - we restrict the super-resolution evaluation to two bias correction models and two baselines. We select UNet𝜆=2 and UNet Multi FT, which achieve the lowest average wet-pixel MSE in the bias correction evaluation (0.066 and 0.065 respectively) and represent two distinct training strategies. We detail the full list of hyperparameters for these final models in the Supplementary Material, Section S3. As baselines, we include B0 (raw AIFS) and B4 (per-pixel multivariate linear regression). We generate 𝑀 = 12 ensemble members for each forecast initialization. The evaluation is performed on approximately one quarter of the test times, while keeping the same lead-time coverage and evaluation protocol (we consider every 4 time steps in the test set, to cover all seasons through the year). Since the same super-resolution
14
Fig. 6: Qualitative bias correction results at 18h (left) and 72h (right) lead time, for random time steps in the test set. Rows show, from top to bottom: raw AIFS input, ground-truth precipitation, and predictions from four of our multi-lead models.
model is used in all pipelines, performance differences reflect only the upstream bias correction; we therefore refer to each pipeline by its bias correction model for brevity.
2) Lead-time dependence of skill
The heatmap shown in Figure 7c summarizes the average performance across all lead times. UNet𝜆=2 achieves the best CRPS (0.434), MSE (3.570), and wet-pixel MSE (10.328). UNet Multi FT achieves the highest CSI (0.324) and avFSS5 (0.450), and notably near-zero amplitude bias (𝑒 𝐴 = −0.010), a substantial improvement over B0 (𝑒 𝐴 = 0.773, over-prediction) and B4 (𝑒 𝐴 = −0.882, underprediction).
Figures 7a and 7b show multi-metric comparisons across lead times (for per-lead values and additional avFSS / Ensemble SAL diagnostics, please see the Supplementary Material, Section S2). Both SwAIther-Precip configurations dominate all metrics at every lead, with UNet𝜆=2 achieving the best CRPS throughout (0.373 at 6 h to 0.532 at 6 d, a 51% CRPS reduction over B0 at short range). Performance is remarkably stable from 6 h through 2 d and degrades gradually beyond, with all models converging at 6 d as predictability decreases. UNet Multi FT trails slightly in CRPS at extended ranges but maintains competitive CSI and avFSS5 , and provides better amplitude calibration overall.
The raw AIFS baseline (B0) unsurprisingly performs worst across all metrics (CRPS 0.835), demonstrating that applying super-resolution to uncorrected AIFS output amplifies rather than compensates for biases. B4 improves upon B0 but remains substantially behind the UNet-based pipelines. The structure component 𝑒𝑆 remains positive for all models (0.473–1.102), reflecting a common tendency of generative models to produce slightly smoother precipitation objects than observed.
Figure 8 summarises PIT diagnostics aggregated over ∼248 × 106 pixel–time samples. Both neural models are substantially better calibrated (KL = 0.182 and 0.208) than AIFS (0.563) and competitive with B4 (0.176). Calibration remains stable up to 3 days before increasing moderately (panel b), while AIFS degrades rapidly beyond 24 h. Stratification by intensity (panel c) reveals that heavy precipitation (≥ 1 mm) is the most challenging regime (KL ≈ 1.7–
1) Performance averaged across lead times
15 CRPS
avFSS
2d
MAE_wet
0.366
0.478
MAE_wet
0.590
0.570
0.482
avFSS MAE_wet
0.457
0.698
CRPS
MAE_wet
avFSS
MAE_wet
swAIther (UNet = 2) swAIther (Multi FT)
CRPS
AIFS raw
avFSS
1.550
B4 baseline
0.465
CRPS
MAE_wet
5d
UNet Single 6h
avFSS
0.430
0.395
CSI
18h
CSI MAE_wet
swAIther (UNet = 2) swAIther (Multi FT)
outward = better on all axes (inverted: CRPS, MAE_wet)
CRPS
AIFS raw B4 baseline
avFSS
4d
UNet Single 3d
UNet Single 6d
outward = better on all axes (inverted: CRPS, MAE_wet)
(a) Short-range horizons (6h – 2 days)
(b) Extended-range horizons (3 days – 6 days)
B0: AIFS (raw)
0.835
5.865
13.637
0.258
0.371
+1.102
+0.773
0.285
B4: Pixel LR (multivar)
0.488
4.935
13.469
0.270
0.407
+0.798
-0.882
0.254
swAIther-Cast (UNet =2)
0.434
3.570
10.328
0.322
0.449
+0.473
-0.313
0.211
swAIther-Cast (Multi FT)
0.469
3.861
10.560
0.324
0.450
+0.704
-0.010
0.229
CRPS
MSE
MSEwet
CSI
avFSS
eS
eA
eL
Better
0.333 0.303
avFSS
CSI
24h
CSI
1.424
0.8251.676 0.244 0.360
CRPS
3d
1.299
0.274
0.432
CSI
MAE_wet
CSI
0.328
1.477
CRPS 0.442
0.348
1.333
0.7021.622 0.287 0.407
avFSS
6d
CSI
0.308
CRPS
6h
1.188
CSI
Worse
(colors normalized per column; for eS and eA, closer to 0 is better)
(c) Average metrics across all lead times
Fig. 7: Full downscaling pipeline evaluation. (a,b) Multi-metric radar charts across forecast horizons. Each quadrant corresponds to a lead time, with four metrics per quadrant: CRPS, MAE (wet pixels), CSI, and avFSS5 . Within each quadrant, metrics are arranged clockwise in this order starting from the quadrant’s left boundary, so that each spoke belongs unambiguously to the quadrant on its right. Outward displacement indicates better performance; single-horizon UNet models appear as filled wedges at their respective lead times. (c) Average metrics across all lead times (colors normalized per column; for eS and eA, closer to zero is better).
3.0); UNet𝜆=2 achieves the best calibration among biascorrected models across all intensity bins.
training provides a regularization benefit that is increasingly important at longer horizons.
3) Multi-lead training versus single-lead specialization
4) Spectral fidelity of the downscaled forecasts
Single-horizon UNet models appear as filled wedges in Figure 7. At 6h, the specialist closely matches SwAItherPrecip (CRPS 0.366 vs. 0.373). At 3d, it falls slightly behind in CSI and avFSS5 . At 6d, the single-lead model degrades substantially (CRPS 0.604 vs. 0.532 for UNet𝜆=2 , MSE wet 14.506 vs. 11.889), confirming that multi-lead
Figures 9 and 10 present the spectral diagnostics across three forecast horizons. Both SwAIther-Precip models achieve spectral ratios close to unity at large scales (0.85– 0.93) and small scales (0.88–0.98), substantially outperforming the baselines: B0 systematically exceeds 1 (reaching 1.7–1.9 at small scales, reflecting AIFS over-prediction of wet extent), while B4 collapses below 0.3 at all scales,
16
PSD ratio (pred / obs)
Spectral Ratio
Fig. 8: PIT calibration diagnostics for both SwAIther-Precip variants and the considered downscaling baselines. (a) Global PIT histograms with KL divergence from a uniform distribution. (b) KL divergence as a function of lead time. (c) KL divergence stratified by precipitation intensity. 1.8 1.6 1.4 1.2 1.0 0.8 0.6 0.4 0.2 0.0
6h large
1000
3days
meso
100
small
10
large
1000
100
6h
6days
meso
small
10
large
1000
meso
100
3days
small
10
6days
PSD [log(P + )]
Co-masked Spectra
104 103 102 101 100 10 1 10 2 1000
100
Wavelength (km)
10 UNet_lam2 MultiFT2
1000 B0: AIFS B4: Pixel LR, MultiVar
100
Wavelength (km) UNet Single 6h UNet Single 3days
10 UNet Single 6days Observed (median)
1000
100
Wavelength (km)
10
Observed (co-mask envelope) 0.5 threshold
Fig. 9: Spectral analysis of the downscaling pipeline across lead times (6h, 3days, 6days). Top row: spectral ratio (predicted/observed PSD) computed on co-masked log-precipitation fields. A ratio of 1 indicates perfect spectral fidelity; the dotted line marks the 0.5 effective-resolution threshold. Shaded bands denote the large-scale (100–600km), mesoscale (20–100km), and small-scale (2–20km) wavelength ranges. Bottom row: radially averaged power spectral densities of log-transformed precipitation, computed within the common wet area where both prediction and truth exceed 0.1 mm/6h (co-masking). The black line shows the median observed spectrum across models, with the gray envelope indicating the range due to model-dependent co-mask differences.
producing fields with less than a third of the observed variance. Both SwAIther-Precip models maintain stable characteristics from 6 h to 3 d, with moderate large-scale degradation at 6 d, and achieve an effective resolution of ∼4 km on the 1 km grid, consistent with state-of-the-art NWP
models (Klaver et al. 2020; Skamarock 2004). Singlehorizon models perform comparably or worse spectrally (UNet Single 6d drops to 0.55 at small scales), further supporting the multi-lead regularization benefit.
17 Large (100 600 km)
6h
Meso (20 100 km)
Small (2 20 km)
3 days
6 days
eff. res.
eff. res.
eff. res.
B0: AIFS
4 km
4 km
4 km
B4: Pixel LR
>1000 km
>1000 km
>1000 km
swAIther-Cast (UNet =2)
4 km
4 km
2 km
swAIther-Cast (Multi FT)
4 km
4 km
4 km
UNet Single
88 km
4 km
313 km
0.00 0.25 0.50 0.75 1.00 1.25 1.50 1.75 2.00 0.00 0.25 0.50 0.75 1.00 1.25 1.50 1.75 2.00 0.00 0.25 0.50 0.75 1.00 1.25 1.50 1.75 2.00 Spectral ratio (pred / obs) Spectral ratio (pred / obs) Spectral ratio (pred / obs)
Fig. 10: Band-averaged spectral ratios (predicted/observed PSD) at three forecast horizons. Each dot represents the ratio averaged over a wavelength band: large scale (100–600km, diamond), mesoscale (20–100km, square), and small scale (2–20km, circle). The dashed line marks perfect spectral fidelity (ratio=1). Effective resolution (km) is shown on the right for each horizon. The SwAIther-Precip models maintain ratios close to 1 across scales, with a consistent mesoscale deficit. B0 (AIFS) systematically exceeds 1 (over-dispersion), while B4 collapses below 0.3 (severe under-dispersion).
The most consistent spectral deficiency is a variance deficit at mesoscales (20–100 km), with band ratios of 0.63–0.79, visible in Figure 10 as the mesoscale marker pulling left of the large- and small-scale ones. This deficit, and its origin in the deterministic bias correction, is analyzed in Section b.
5) Ensemble characteristics and sample quality Figure 11 presents ensemble members for a strong winter precipitation event (2019-02-03, 2-day lead). For both pipelines, ensemble members share a consistent largescale structure while exhibiting meaningful diversity in fine-scale details - edges of the precipitation region, isolated features, and the shape of the high-intensity core vary across members, reflecting genuine forecast uncertainty at convective scales. The quantile summaries (Q05, median, Q95) reveal well-calibrated ensemble spread, though both pipelines seem to slightly underestimate peak intensities compared to the truth, an attenuation inherited from the deterministic bias correction. UNet Multi FT produces noticeably higher peak intensities than UNet𝜆=2 , consistent with its near-zero amplitude bias (𝑒 𝐴 = −0.010 vs. −0.313) and higher mesoscale spectral ratio (0.79 vs. 0.69 at 3 days). The multi-step fine-tuning partially mitigates the variance compression inherent in the deterministic bias correction, yielding more intense fields at the cost of slightly higher MSE (3.861 vs. 3.570).
Figure 12 shows ensemble-median predictions across all lead times. Both pipelines maintain remarkably stable spatial structure from 6h through 3 d, with gradual degradation at 4–6 d reflecting increasing synoptic-scale uncertainty. Despite progressive smoothing of the bias-corrected input at longer horizons, the diffusion model continues to generate realistic fine-scale texture at all lead times, consistent with the stable small-scale spectral ratios reported above. Additional ensemble-median predictions across all lead times for a particularly heavy precipitation event from November 12th to November 15th , 2023 can be observed in the Suspplemental Material, Section S5 6. Discussion and conclusions a. Summary of findings The results presented in Sections 5a–5b support the three main claims of this work. The two-step decomposition is effective. The full pipeline achieves a 48% CRPS reduction over raw AIFS and 11% over the strongest classical baseline, with well-calibrated ensemble spread and skill across all metrics (Figure 7). The Strategy 2 proxy ablation (Section 4) shows substantially degraded performance when bias correction and superresolution are collapsed into a single step, suggesting that the resolution gap and systematic biases in AIFS precipitation benefit from explicit intermediate correction. Lead-time awareness is both necessary and efficient. A single lead-conditioned model matches or outperforms
18
Fig. 11: High-resolution precipitation samples generated by the conditional diffusion model (Step 2) for our two SwAIther-Precip variants (using UNet𝜆=2 and Multi FT bias-correction backbones) at 2-day lead time, for a random time step in the test set. Left column: high-resolution ground truth and raw AIFS input. Per model: the deterministic bias-corrected field used as conditioning input, four ensemble members (out of the existing 12 members), and summary statistics (Q05, median, Q95) computed over all 20 members. Individual samples exhibit realistic fine-scale variability while the ensemble median recovers the broad spatial structure of the ground truth.
horizon-specific specialists at all lead times, with the advantage growing at longer horizons (13% CRPS improvement at 6 days; Figure 7). Spectral quality remains stable across horizons, with effective resolution at ∼4 km from 6 h through 6 days (Figure 10). This is operationally attractive, as a single model replaces a suite of per-horizon specialists. The end-to-end pipeline produces km-scale, medium-range forecasts with realistic spatial structure. Spectral ratios remain close to unity at both large and small scales across all horizons (Figure 10), and the ∼4 km effective resolution on a 1 km grid is consistent with state-of-the-art NWP models (Klaver et al. 2020; Skamarock 2004). Ensemble members exhibit realistic fine-scale texture with meaningful inter-member diversity (Figure 11). b. Limitations and future work Several limitations of the current framework suggest directions for future work. The main limitation is a recurring trade-off between spatial bias correction and variance preservation. The deterministic bias correction excels at correcting spatial placement and extent of precipitation, reducing AIFS’s wet fraction from ∼58% to a realistic ∼28%, but tends to underestimate peak intensities (negative SAL-A values; Figure 3) and exhibits a persistent mesoscale variance deficit (band ratios of 0.63–0.79; Figure 10). This is a direct consequence of the Huber loss, which regresses toward the conditional mean at scales where precipitation is genuinely stochastic. The CorrDiff super-resolution partially compensates
at small scales (2–20 km) but cannot recover mesoscale variance lost in the intermediate representation. UNet Multi FT—fine-tuned with a multi-step loss—partially mitigates this effect, achieving the best amplitude calibration (𝑒 𝐴 = −0.010 at the super-resolution level) and the highest mesoscale spectral ratios (0.79 at 3 days), and visibly capturing peak intensities better than UNet𝜆=2 in qualitative examples (Figures 6 and 11). The intensity attenuation is particularly concerning for extreme events, which are precisely the cases where high-resolution probabilistic forecasts are most valuable. Promising mitigation directions include spectral regularization terms in the bias correction loss, stochastic bias correction (e.g., conditional diffusion or VAE) that can represent the multi-modal nature of mesoscale precipitation, multi-step or spectral-aware losses building on the Multi FT result, and training the CorrDiff super-resolution to target a broader spectral range. Future work should also specifically evaluate the pipeline’s performance on heavy precipitation events and explore loss functions that preserve tail behavior. Another scope limitation is the lack of temporal consistency: the super-resolution samples are independent when conditioned on the bias-correction outputs. Extending the diffusion model to generate temporally consistent trajectories is an important direction for applications requiring consistent precipitation time-series or accumulations. Our comparison with single-step direct downscaling used only the regression component of CorrDiff as a proxy, since running the complete diffusion model on all input channels was computationally prohibitive. A complete end-to-end
Fig. 12: High-resolution precipitation median samples across all considered lead times (6h to 6d), for our two SwAIther-Precip variants (using UNet𝜆=2 and Multi FT bias-correction backbones), for a random time step in the test set. Top left: high-resolution ground truth. First row: raw AIFS input at each horizon. Per model: the deterministic bias-corrected field and the ensemble median of the conditional diffusion model (Step 2).
19
20 pipeline for the “direct task” (Strategy 2) would provide a more definitive comparison. As our framework is trained over Switzerland using AIFS and CombiPrecip, performance may depend on domain characteristics (orographic complexity, convective regime) and the bias structure of the driving model. Evaluation on other domains and with other global models would help assess generalizability.
Acknowledgments. TB acknowledge support from the Swiss National Science Foundation (SNSF) under Grant No. 10001754 (“RobustSR” project). DA, FA, TL, KV, EK, and TB acknowledge support from the Swiss Data Science Center (SDSC) End-User Innovation Project under grant CI24-04: NWF4CH - Democratizing Neural Weather Forecasting for Switzerland : An Open Platform Approach. We thank Max Defez, Shivanshi Asthana, Milton Gomez, David Leutwyler, Mary McGlohon, Petar Stamenkovic for advice that helped develop SwAIther-Precip. Data availability statement. In order to recreate the AIFS predictions historical dataset, recent ECMWF IFS step-0 forecast fields are available via ECMWF Open Data (https://data.ecmwf.int/forecasts/), but historical operational IFS initial conditions are only available through ECMWF archive access subject to ECMWF member access conditions. CombiPrecip data are publicly available (https://opendatadocs.meteoswiss.ch/). APPENDIX A Appendix: Model architecture details We present in Figure A1 the detailed architecture of the UNet model we use in both steps of the SwAIther framework, as well as architectural details for both steps in Figure A2. For a detailed list of hyperparameters, and discussions regarding choices the bias correction output clamp value and the noise-level embedding of the Step 2 diffusion model, please see the Supplementary Material (Section S3). References Abdi, D., I. Jankov, P. Madden, V. Vargas, T. A. Smith, S. Frolov, M. Flora, and C. Potvin, 2026: Hrrrcast: A data-driven emulator for regional weather forecasting at convection-allowing scales. Artificial Intelligence for the Earth Systems, 5 (2), 250 061, https://doi.org/ 10.1175/AIES-D-25-0061.1. Adamov, S., and Coauthors, 2025: Building machine learning limited area models: Kilometer-scale weather forecasting in realistic settings. URL https://arxiv.org/abs/2504.09340, 2504.09340. Andrae, M., T. Landelius, J. Oskarsson, and F. Lindsten, 2024: Continuous ensemble weather forecasting with diffusion models. arXiv preprint arXiv:2410.05431. Antonio, B., A. T. T. McRae, D. MacLeod, F. C. Cooper, J. Marsham, L. Aitchison, T. N. Palmer, and P. A. G. Watson, 2025:
Postprocessing east african rainfall forecasts using a generative machine learning model. Journal of Advances in Modeling Earth Systems, 17 (3), e2024MS004 796, https://doi.org/https://doi.org/10. 1029/2024MS004796, https://agupubs.onlinelibrary.wiley.com/doi/ pdf/10.1029/2024MS004796. Arai, R., T. Sato, and M. Imamura, 2025: Enhancing the spatial resolution of medium-range precipitation forecasts using super-resolution neural networks. Weather and Forecasting, 40 (12), 2561 – 2578, https://doi.org/10.1175/WAF-D-24-0217.1. Badrinath, A., L. D. Monache, N. Hayatbini, W. Chapman, F. Cannon, and M. Ralph, 2023: Improving precipitation forecasts with convolutional neural networks. Weather and Forecasting, 38 (2), 291 – 306, https://doi.org/10.1175/WAF-D-22-0002.1. Bengio, Y., J. Louradour, R. Collobert, and J. Weston, 2009: Curriculum learning. Proceedings of the 26th annual international conference on machine learning, 41–48. Bi, K., L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian, 2023: Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619 (7970), 533–538. Chen, H., T. Han, J. Zhang, S. Guo, and L. Bai, 2025: Stcast: Adaptive boundary alignment for global and regional weather forecasting. arXiv preprint arXiv:2509.25210. Dawid, A. P., 1984: Present position and potential developments: Some personal views statistical theory the prequential approach. Journal of the Royal Statistical Society: Series A (General), 147 (2), 278–290. Espeholt, L., and Coauthors, 2022: Deep learning for twelve hour precipitation forecasts. Nature communications, 13 (1), 5145. Gao, Y., and Coauthors, 2025: OneForecast: A universal framework for global and regional weather forecasting. Proceedings of the 42nd International Conference on Machine Learning, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu, Eds., PMLR, Proceedings of Machine Learning Research, Vol. 267, 18 658–18 697. Gneiting, T., and A. E. Raftery, 2007: Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102 (477), 359–378. Harris, L., A. T. T. McRae, M. Chantry, P. D. Dueben, and T. N. Palmer, 2022: A generative deep learning approach to stochastic downscaling of precipitation forecasts. Journal of Advances in Modeling Earth Systems, 14 (10), e2022MS003 120, https://doi.org/https://doi.org/10. 1029/2022MS003120, https://agupubs.onlinelibrary.wiley.com/doi/ pdf/10.1029/2022MS003120. He, K., X. Zhang, S. Ren, and J. Sun, 2016: Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778. Hersbach, H., 2000: Decomposition of the continuous ranked probability score for ensemble prediction systems. Weather and Forecasting, 15 (5), 559–570. Hessami, M., P. Gachon, T. B. Ouarda, and A. St-Hilaire, 2008: Automated regression-based statistical downscaling tool. Environmental modelling & software, 23 (6), 813–834. Hobeichi, S., N. Nishant, Y. Shao, G. Abramowitz, A. Pitman, S. Sherwood, C. Bishop, and S. Green, 2023: Using machine learning to cut the cost of dynamical downscaling. Earth’s Future, 11 (3), e2022EF003 291, https://doi.org/https://doi.org/10.1029/
4×[H×H×64] blocks
4×
[ H × H × 128 2 2 blocks
DownBlock
[
]
(2 + 5) ×
UpBlock
]
5×
H × H × 128 4 4 blocks
[
self-attn post-attn
×5 skips
Bottleneck
H × H × 128 4 4 blocks
DownBlock
4×
]
×5 skips
×5 skips (copy + concat each)
[
]
5×[H×H×64] blocks
FiLM shift
Fig. A1: UNet Architecture, shared in both SwAIther step.
Note: the long-range encoder→decoder skips drawn above are unscaled feature concatenations, √ distinct from the per-block 1/ 2-rescaled internal ResNet residual.
DownBlock: UNetBlock with stride-2 conv0 (and a strided 1×1 conv on the shortcut). UpBlock: UNetBlock preceded by 2× nearest-neighbour upsampling (and an upsampled 1×1 conv on the shortcut). √ Self-attn: UNetBlock followed by a multi-head self-attention layer, residually combined and rescaled by 1/ 2. Post-attn: a plain UNetBlock placed immediately after Self-attn to mix the attended features. No encoder-skip concat.
(normi = GroupNorm+SiLU; convi = 3×3 Conv2D , with conv1 zero-initialised; affine = Linear from cond. emb. → per-channel shift.) When input and output channel counts differ, the shortcut passes through a 1×1 conv; otherwise it is the identity.
√ UNetBlock: norm0 → conv0 → affine → norm1 → conv1, plus a 1/ 2-rescaled internal ResNet shortcut from input to output.
Architecture of the blocks:
Output
H×W×Cout
Conv2D
Decoder (up-sampling)
H × H × 128 2 2 blocks
UpBlock
Each cuboid = one residual block. Front face = spatial dim H×W ; depth = channel dim C. Each level pushes 5 skips onto a shared stack (initial conv/down + 4 blocks). Each decoder block (except self-attn and post-attn) pops one skip and concats it with its running input, paired one-to-one in LIFO order.
Encoder (down-sampling)
H×W×Cin
Conv2D
Input
21
22
(a) Step 1 – Bias correction: lead-time-aware FiLM + clamped residual U-Net (3 levels) Lead time ℓ
Learned embedding
2-layer MLP
γ ℓ , βℓ
integer 0–23
ℓ 7→ zℓ ∈ Remb_dim
emb_dim → mlp_hidden → 2Cin
per-channel scale & shift
FiLM modulation x · (1 + γ ℓ ) + β ℓ
U-Net ϕ
γ∆
predicts residual
+
∆ = ϕ(FiLM(x, ℓ))
AIFS input channels [xt,ℓ ]
Bias-corrected Precip [e xt,ℓ ]
(prec) Anchor p̄0 = x̄t,ℓ
precip channel after FiLM residual / skip from input
(b) Step 2 – Generative super-resolution: CorrDiff regression + conditional diffusion
Bias-corrected precip + altitude
Regression U-Net rψ
2 ch., H×W
H ×W → sH ×sW
(with upsampling)
b reg y
sH ×sW
Condition on
Noise zτ sH ×sW
Diffusion U-Net dω (conditioned + noise embed.) residual generator
b y
reg
∆y
+
e ,x
High-res precip bt,ℓ y sH ×sW
diff
sH ×sW
Zero embedding 0 ∈ RCemb
no noise-level cond. Diffusion U-Net (EDMPrecondSuperResolution): same building block as the UNet used for the bias correction but with a noise-timestep embedding fed to all blocks. Self-attention placed at the deepest level
FiLM block
UNet model
Input
Output
Conditioning input
+ Element-wise sum
Fig. A2: Detailed architecture for Step 1 (Bias correction) and Step 2 (Super-resolution). The U-Net architecture used within the different steps is detailed in Figure A1. Note that the Super-resolution architecture (b) is mostly taken from the CorrDiff implementation.
2022EF003291, https://agupubs.onlinelibrary.wiley.com/doi/pdf/10. 1029/2022EF003291.
Lang, S., and Coauthors, 2024: Aifs–ecmwf’s data-driven forecasting system. arXiv preprint arXiv:2406.01465.
Klaver, R., R. Haarsma, P. L. Vidale, and W. Hazeleger, 2020: Effective resolution in high resolution global atmospheric models for climate studies. Atmospheric Science Letters, 21 (4), e952.
Larsson, E., J. Oskarsson, T. Landelius, and F. Lindsten, 2025: Crps-lam: Regional ensemble weather forecasting from matching marginals. URL https://arxiv.org/abs/2510.09484, 2510.09484.
Lam, R., and Coauthors, 2023: Learning skillful medium-range global weather forecasting. Science, 382 (6677), 1416–1421, https://doi.org/ 10.1126/science.adi2336, https://www.science.org/doi/pdf/10.1126/ science.adi2336.
Leinonen, J., U. Hamann, D. Nerini, U. Germann, and G. Franch, 2023: Latent diffusion models for generative precipitation nowcasting with accurate uncertainty quantification. arXiv preprint arXiv:2304.12891.
23 Li, B., Z. Zhu, X. Zhong, R. Tan, Y. Wang, W. Lan, and H. Li, 2025: One-kilometer resolution forecasts of hourly precipitation over china using machine learning models. Atmospheric Science Letters, 26 (3), e1297. Mardani, M., and Coauthors, 2025: Residual corrective diffusion modeling for km-scale atmospheric downscaling. Communications Earth & Environment, 6 (1), 124. Matheson, J. E., and R. L. Winkler, 1976: Scoring rules for continuous probability distributions. Management Science, 22 (10), 1087–1096. Moldovan, G., and Coauthors, 2025: An update to ecmwf’s machine-learned weather forecast model aifs. arXiv preprint arXiv:2509.18994. Molinaro, R., N. Siegenheim, H. Martin, M. Frey, N. Poulsen, P. Seitz, and M. V. Gabler, 2026: Universal diffusion-based probabilistic downscaling. URL https://arxiv.org/abs/2602.11893, 2602.11893. Necker, T., J. Hackenbruch, D. Reinert, C. Gebhardt, E. Bauernschubert, and R. Potthast, 2024: A consistent definition of the fractions skill score for ensemble forecasts. Quarterly Journal of the Royal Meteorological Society, https://doi.org/10.1002/qj.4XXX. Nipen, T. N., and Coauthors, 2025: Regional data-driven weather modeling with a global stretched-grid. Artificial Intelligence for the Earth Systems, https://doi.org/10.1175/AIES-D-25-0001.1. Pathak, J., and Coauthors, 2026: Kilometer-scale convection-allowing model emulation using generative diffusion modeling. Science Advances, 12 (5), eadv0423, https://doi.org/10.1126/sciadv.adv0423, https://www.science.org/doi/pdf/10.1126/sciadv.adv0423. Perez, E., F. Strub, H. De Vries, V. Dumoulin, and A. Courville, 2018: Film: Visual reasoning with a general conditioning layer. Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Price, I., and S. Rasp, 2022: Increasing the accuracy and resolution of precipitation forecasts using deep generative models. Proceedings of The 25th International Conference on Artificial Intelligence and Statistics, G. Camps-Valls, F. J. R. Ruiz, and I. Valera, Eds., PMLR, Proceedings of Machine Learning Research, Vol. 151, 10 555–10 571. Pulkkinen, S., D. Nerini, A. A. Pérez Hortal, C. Velasco-Forero, A. Seed, U. Germann, and L. Foresti, 2019: Pysteps: An open-source python library for probabilistic precipitation nowcasting (v1. 0). Geoscientific Model Development, 12 (10), 4185–4219. Radanovics, S., J.-P. Vidal, and E. Sauquet, 2018: Verification of spatial precipitation forecasts using the fractions skill score: the ensemble case. Weather and Forecasting, 33 (3), 655–670. Rasp, S., and Coauthors, 2024: Weatherbench 2: A benchmark for the next generation of data-driven global weather models. Journal of Advances in Modeling Earth Systems, 16 (6), e2023MS004 019, https://doi.org/https://doi.org/10.1029/ 2023MS004019, https://agupubs.onlinelibrary.wiley.com/doi/pdf/ 10.1029/2023MS004019. Ravuri, S., and Coauthors, 2021: Skilful precipitation nowcasting using deep generative models of radar. Nature, 597 (7878), 672–677. Roberts, N. M., and H. W. Lean, 2008: Scale-selective verification of rainfall accumulations from high-resolution forecasts of convective events. Monthly Weather Review, 136 (1), 78–97. Rojas-Campos, A., M. Langguth, M. Wittenbrink, and G. Pipa, 2023: Deep learning models for generation of precipitation maps based
on numerical weather prediction. Geoscientific Model Development, 16 (5), 1467–1480, https://doi.org/10.5194/gmd-16-1467-2023. Sha, Y., and Coauthors, 2026: Ai-based regional emulation for kilometer-scale dynamical downscaling. arXiv preprint arXiv:2602.18646. Siddiqui, S. A., J. Kossaifi, B. Bonev, C. Choy, J. Kautz, D. Krueger, and K. Azizzadenesheli, 2024: Exploring the design space of deep-learning-based weather forecasting systems. arXiv preprint arXiv:2410.07472. Sideris, I., M. Gabella, M. Sassi, and U. Germann, 2014: The combiprecip experience: development and operation of a real-time radarraingauge combination scheme in switzerland. 2014 international weather radar and hydrology symposium, 1–10. Sinclair, S., and G. Pegram, 2005: Empirical mode decomposition in 2-d space and time: a tool for space-time rainfall analysis and nowcasting. Hydrology and Earth System Sciences, 9 (3), 127–137. Skamarock, W. C., 2004: Evaluating mesoscale nwp models using kinetic energy spectra. Monthly weather review, 132 (12), 3019–3032. Sønderby, C. K., and Coauthors, 2020: Metnet: A neural weather model for precipitation forecasting. arXiv preprint arXiv:2003.12140. Sun, H., H. Jing, Z. Dai, S. Xiao, W. Xue, J. Sun, and Q. Lu, 2026: China regional 3 km downscaling based on residual corrective diffusion model. EGUsphere, 2026, 1–26, https://doi.org/ 10.5194/egusphere-2026-822. Wan, Z. Y., R. Baptista, Y. fan Chen, J. Anderson, A. Boral, F. Sha, and L. Zepeda-Núñez, 2023: Debias coarsely, sample conditionally: Statistical downscaling through optimal transport and probabilistic diffusion models. URL https://arxiv.org/abs/2305.15618, 2305.15618. Wernli, H., M. Paulat, M. Hagen, and C. Frei, 2008: Sal—a novel quality measure for the verification of quantitative precipitation forecasts. Monthly Weather Review, 136 (11), 4470–4487. Wilby, R. L., C. W. Dawson, and E. M. Barrow, 2002: Sdsm—a decision support tool for the assessment of regional climate change impacts. Environmental Modelling & Software, 17 (2), 145–157. Xu, H., Y. Zhao, D. Zhao, Y. Duan, and X. Xu, 2024: Improvement of disastrous extreme precipitation forecasting in north china by panguweather ai-driven regional wrf model. Environmental Research Letters, 19 (5), 054 051. Xu, P., and Coauthors, 2025: An artificial intelligence-based limited area model for forecasting of surface meteorological variables. Communications Earth & Environment, 6 (1), 372.
24 APPENDIX S Supplementary Material S1. AIFS-Single-1.0: Input/Output Variable List We present in Table S1 and Table S2 the list of input and output fields for the AIFS-single-1.0 model we use to create a dataset of precipitation forecasts, input to our downscaling framework. Table S1: Variables of interest for the AIFS-single-1.0 model. Variable name
Description
Units
z t u v w q sp msl skt sst 2t 2d 10u 10v 100u 100v tcw swvl1 swvl2 stl1 stl2 tp cp sf rowe ssrd strd tcc hcc mcc lcc lsm sdor slor lat lon insol tod doy
Geopotential (pressure levels) Air temperature (pressure levels) Zonal wind component (pressure levels) Meridional wind component (pressure levels) Vertical velocity (pressure coordinates) Specific humidity (pressure levels) Surface pressure Mean sea-level pressure Skin temperature Sea surface temperature 2 m air temperature 2 m dewpoint temperature 10 m zonal wind component 10 m meridional wind component 100 m zonal wind component 100 m meridional wind component Total column water Soil moisture layer 1 (volumetric) Soil moisture layer 2 (volumetric) Soil temperature layer 1 Soil temperature layer 2 Total precipitation Convective precipitation Snowfall water equivalent Runoff water equivalent Surface solar radiation downwards Surface thermal radiation downwards Total cloud cover High cloud cover Medium cloud cover Low cloud cover Land-sea mask Standard deviation of sub-grid orography Slope of sub-grid orography Latitude Longitude Insolation forcing Time of day (cyclical encoding) Day of year (cyclical encoding)
m2 s −2 K m s −1 m s −1 Pa s −1 kg kg−1 Pa Pa K K K K m s −1 m s −1 m s −1 m s −1 kg m −2 m3 m −3 m3 m −3 K K kg m −2 kg m −2 kg m −2 kg m −2 J m −2 J m −2 0–1 (fraction) 0–1 (fraction) 0–1 (fraction) 0–1 (fraction) 0–1 (fraction) m — degrees degrees — — —
25 Table S2: Full list of input and output fields for the AIFS-single-1.0 model. Field
Level type
Input/Output
Geopotential, horizontal and vertical wind components, specific humidity, temperature
Pressure level: 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 850, 925, 1000
Both
Surface pressure, mean sea-level pressure, skin temperature, 2 m temperature, 2 m dewpoint temperature, 10 m horizontal wind components, total column water
Surface
Both
Soil moisture and soil temperature (layers 1 and 2)
Surface
Both
100m horizontal wind components, solar radiation (Surface short-wave (solar) radiation downwards and Surface long-wave (thermal) radiation downwards), cloud variables (tcc, hcc, mcc, lcc), runoff and snow fall
Surface
Output
Total precipitation, convective precipitation
Surface
Output
Land-sea mask, orography, standard deviation of subgrid orography, slope of sub-scale orography, insolation, latitude/longitude, time of day/day of year
Surface
Input
S2. Detailed results by Lead Time a. Bias correction metrics by lead Time The bias correction performance results were presented in the main text for a limited choice of lead times and neighborhood sizes regarding the SAL and FSS metric. To complement these results, Figures S1 and S2 present detailed results for the SAL and FSS metrics, across all lead times. In particular, S, A and L metrics are shown in Figure S1, and FSS metrics are given for three values of neighborhood size (𝑛 = 1, 𝑛 = 5, and 𝑛 = 17) in Figure S2.
Fig. S1: Lead-time evolution of SAL metrics for the bias correction task (Step 1). S, A and L metrics are shown across all leads. Results are shown for multi-lead UNet and ViT variants, all baselines, and single-lead specialist models evaluated at their respective target leads.
b. Downscaling metrics by lead Time Table S3 shows detailed per lead detailed results for the downscaling pipelines we consider in the main manuscript. In particular, it shows CRPS, CSI, member-averaged FSS5 , and MSE (wet pixels) by lead time across our models and baselines (including single lead time models). Figure S3 visualizes the evolution of CRPS, CSI and MSE (wet pixels) with lead time, using line plots. Figures S5 and S4 visualize the evolution of avFSS for different neighborhood choices and eS, eA and eL within the ensemble SAL metric with lead time, using line plots.
26
Fig. S2: Lead-time evolution of FSS metrics for the bias correction task (Step 1). Three different values of neighborhoods are considered here, across all leads, to complement the results shown in the main text. Results are shown for multi-lead UNet and ViT variants, all baselines, and single-lead specialist models evaluated at their respective target leads.
Fig. S3: Lead-time evolution of key downscaling metrics for the full two-step pipeline. Results are shown for multi-lead UNet variants, baselines (B0, B4), and single-lead specialist models evaluated at their respective target leads.
Fig. S4: Lead-time evolution of SAL metrics for the full two-step pipeline. Results are shown for multi-lead UNet variants, baselines (B0, B4), and single-lead specialist models evaluated at their respective target leads.
27 Table S3: Downscaling (Step 2): CRPS, CSI, member-averaged FSS5 , and MSE (wet pixels) by lead time across our models and baselines (including single lead time models). Model
6h
12h
18h
1d
2d
3d
4d
5d
6d
Avg
B0: AIFS B4: Pixel LR, MultiVar
0.770 0.482
0.772 0.481
0.793 0.481
0.806 0.482
0.814 0.484
0.841 0.487
0.882 0.493
0.888 0.497
0.953 0.506
0.835 0.488
SwAIther-Precip (UNet𝜆=2 ) SwAIther-Precip (UNet Multi FT)
0.373 0.384
0.383 0.407
0.388 0.417
0.392 0.416
0.417 0.448
0.450 0.497
0.479 0.543
0.494 0.539
0.532 0.569
0.434 0.469
UNet Single 6h UNet Single 3d UNet Single 6d
0.366 — —
— — —
— — —
— — —
— — —
— 0.442 —
— — —
— — —
— — 0.604
— — —
B0: AIFS B4: Pixel LR, MultiVar
0.267 0.288
0.272 0.290
0.271 0.289
0.270 0.290
0.267 0.290
0.261 0.274
0.249 0.257
0.241 0.238
0.222 0.215
0.258 0.270
SwAIther-Precip (UNet𝜆=2 ) SwAIther-Precip (UNet Multi FT)
0.344 0.341
0.339 0.342
0.341 0.335
0.346 0.348
0.340 0.343
0.333 0.332
0.308 0.311
0.287 0.294
0.262 0.268
0.322 0.324
UNet Single 6h UNet Single 3d UNet Single 6d
0.337 — —
— — —
— — —
— — —
— — —
— 0.319 —
— — —
— — —
— — 0.271
— — —
B0: AIFS B4: Pixel LR, MultiVar
0.383 0.424
0.389 0.424
0.386 0.419
0.386 0.423
0.382 0.430
0.378 0.417
0.362 0.398
0.352 0.376
0.326 0.353
0.372 0.407
SwAIther-Precip (UNet𝜆=2 ) SwAIther-Precip (UNet Multi FT)
0.472 0.467
0.470 0.474
0.474 0.461
0.482 0.481
0.468 0.477
0.465 0.463
0.435 0.431
0.400 0.413
0.376 0.384
0.449 0.450
UNet Single 6h UNet Single 3d UNet Single 6d
0.469 — —
— — —
— — —
— — —
— — —
— 0.448 —
— — —
— — —
— — 0.388
— — —
B0: AIFS B4: Pixel LR, MultiVar
13.179 13.369
12.955 13.370
13.500 13.371
13.331 13.384
13.101 13.413
13.812 13.463
14.174 13.509
14.940 13.610
13.738 13.726
13.637 13.469
SwAIther-Precip (UNet𝜆=2 ) SwAIther-Precip (UNet Multi FT)
9.224 9.032
9.437 9.608
9.520 9.447
9.578 9.655
10.040 10.338
10.766 11.409
11.087 11.828
11.408 11.890
11.889 11.832
10.328 10.560
UNet Single 6h UNet Single 3d UNet Single 6d
9.347 — —
— — —
— — —
— — —
— — —
— 10.786 —
— — —
— — —
— — 14.506
— — —
CRPS ↓
CSI ↑
FSS5 ↑
MSE (wet pixels) ↓
Fig. S5: Lead-time evolution of avFSS metrics for the full two-step pipeline. Results are shown for multi-lead UNet variants, baselines (B0, B4), and single-lead specialist models evaluated at their respective target leads.
28 S3. Architectural details - hyperparameters a. Detailed hyperparameter lists The choice of hyperparameters for the different models used in the methodology are presented in the following section. Note that we only present the two top performing bias correction models. Regarding the super-resolution step, the same CorrDiff submodels (one regression + one diffusion model) is used on top of all bias cocrrection models tested; the hyperparameters for these two models are also presented here. (Step 1) UNet𝜆=2 hyperparameters 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
# --- Architecture (FiLM bias correction U-Net) --model_channels: 64 channel_mult: [1, 2, 2] attn_resolutions: [4] N_grid_channels: 4 embedding_type: zero clamp_min: 0 clamp_max: 20 # --- FiLM lead-time conditioning --n_leads: 24 emb_dim: 64 mlp_hidden: 256 # --- Loss --huber_delta: 1.0 lambda_time: 2.0 lambda_time_warmup_epochs: 5 # --- Optimization --lr: 2e-4 seed: 42 total_batch_size: 1024 fp_optimizations: amp-bf16
(Step 1) MultiFT hyperparameters 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27
# --- Architecture (FiLM bias correction U-Net) --model_channels: 64 channel_mult: [1, 2, 2] attn_resolutions: [4] N_grid_channels: 4 embedding_type: zero # --- FiLM lead-time conditioning --n_leads: 24 emb_dim: 64 mlp_hidden: 256 # --- Output --clamp_min: 0 clamp_max: 20 # --- Training (constant across stages) --total_batch_size: 1024 optimizer: Adam loss: Huber (delta=1.0) lr_decay_rate: 5e5 fp_optimizations: amp-bf16 seed: 42 # --- Multi-stage fine-tuning schedule --# stage lead_index lr lam_time lam_warmup 1: [0, 7] 2e-4 0.0 -2: [8, 15] 2e-4 0.0 -3: [0, 23] 2e-5 1.0 10 4: [12, 23] 1e-5 0.3 3 5: [16, 23] 2e-5 0.0 3
29 (Step 2) Regression U-Net hyperparameters 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
# --- Architecture (CorrDiff regression U-Net) --model_type: SongUNetPosEmbd model_channels: 64 channel_mult: [1, 2, 2] attn_resolutions: [152] N_grid_channels: 4 gridtype: sinusoidal # default; positional grid embed hr_mean_conditioning: False # --- Training --optimizer: Adam loss: L2 (regression) lr: 2e-4 lr_decay_rate: 5e5 total_batch_size: 2 fp_optimizations: amp-bf16 seed: 42
(Step 2) Diffusion U-Net hyperparameters 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
# --- Architecture (CorrDiff diffusion U-Net, EDM) --model_type: SongUNetPosEmbd model_channels: 64 channel_mult: [1, 2, 2] attn_resolutions: [152] N_grid_channels: 4 gridtype: sinusoidal embedding_type: zero # noise-level conditioning disabled # (departs from standard EDM; see SI) hr_mean_conditioning: True # conditions on regression output # --- Training --optimizer: Adam loss: EDM denoising lr: 2.e-4 lr_decay_rate: 5e5 total_batch_size: 2 fp_optimizations: amp-bf16 seed: 42 # --- Conditioning --regression_checkpoint: Step 2 regression U-Net (above)
b. Sensitivity to the output cap We retain a [0, 20] mm/6 h output cap for the Step 1 bias-correction model. The cap was originally introduced as a numerical safeguard during early training. While the climatological 99th percentile of CombiPrecip wet-pixel intensities reaches ∼ 23 mm/6 h on our domain, we observed that raising the cap to [0, 100] mm/6 h gave only marginal differences at Step 1 (CSI: 0.539 → 0.542; FSS5 : 0.828 → 0.834; MSEwet : 0.066 → 0.071) and a small but consistent degradation in end-to-end skill once composed with the Step 2 generative super-resolution model (CRPS: 0.434 → 0.437; CSI: 0.322 → 0.317; FSS5 : 0.449 → 0.440; MSEwet : 10.33 → 10.51). We hypothesize that the tighter cap acts as a soft regularization on Step 1 outputs, producing a more spatially coherent and climatologically smoother conditioning signal for the Step 2 generative model, which is trained on bilinearly-downscaled CombiPrecip and therefore expects similarly bounded inputs. A more thorough analysis of this interaction is left for future work. c. Sensitivity to the noise-level embedding We tested two configurations for the noise-level embedding of the Step 2 diffusion model (second part of CorrDiff): embedding type=zero (constant zero embedding, no explicit noise-level conditioning) and embedding type=positional (the standard EDM sinusoidal positional embedding of the diffusion time-step). The positional variant achieved lower training loss and produced deterministically sharper samples (MSEwet :
30 10.33 → 10.29; CSI: 0.322 → 0.325), but generated less-dispersed ensembles, resulting in worse probabilistic skill (CRPS: 0.434 → 0.447). This is consistent with a known tension between sharpness and calibration in probabilistic forecasting (Gneiting and Raftery 2007): the simpler zero variant, by averaging implicitly across the diffusion noise schedule, appears to produce ensembles whose spread better matches the actual forecast uncertainty distribution, while the positional variant produces more accurate but underdispersed samples. We retain embedding type=zero as the final configuration based on CRPS, which is the primary probabilistic skill metric in our evaluation framework. S4. Additional Experiments a. Summary This section presents additional diagnostic experiments mentioned in Section 3c (Data Section, Weekly train-test split subsection) of the main text. We address two potential concerns regarding the weekly train–test split: (i) whether temporal autocorrelation leads to an overestimation of model performance, and (ii) whether the model trained on the 06–12 and 18–00 UTC windows can generalize to the held-out windows. Numerical results are reported in Tables S4, S5 and S6. We summarize the three main findings: 1. Temporal autocorrelation is not a significant confound. The field-level autocorrelation function of 6-hour accumulated precipitation over Switzerland drops below 0.2 within 12 hours (see supplementary material, Figure 4), indicating that consecutive samples are largely independent at the pixel level. A neutral-terrain comparison further confirms this: the weekly-trained model evaluated exclusively on 2023 (CRPS = 0.533) performs comparably to a yearly-trained model on the same period (CRPS = 0.507), showing that the weekly split does not inflate performance through leaked autocorrelation (Tables 4, 5 and 6 in the supplementary material). 2. Performance differences across years are driven by distribution shift, not leakage. The seasonal decomposition of the yearly model’s 2023 evaluation reveals a ×13 variation in MSEwet between winter (2.4) and summer (31.5), explaining the apparent degradation when evaluating on the full year 2023 compared to the weekly validation set. 3. Performance on held-out time windows reflects intrinsic difficulty, not a generalization gap. A modelindependent comparison of raw AIFS forecasts against observations shows that the 12–18 UTC window (afternoon convection) is nearly twice as difficult as the 00–06 UTC window (×1.96 MSEwet ratio, Table 7 in the supplementary material). The training set already includes a window of comparable difficulty (18–00 UTC, ×1.85). b. Experiments We define multiple experiment labels used throughout, based on the same architecture for the bias correction model, the UNet𝜆=2 model. For every experiment, we evaluate the full pipeline inference, at the high resolution level. The experiments are as follows: • UNet𝜆=2 : reference model trained on the weekly split, evaluated on its own test set (06–12 and 18–00 UTC windows only). • UNet𝜆=2 , weekly, new Hours: same weekly-trained model, evaluated on the held-out windows (00–06 and 12– 18 UTC combined). • UNet𝜆=2 , weekly new Hours, only 18: same model, evaluated on the 12–18 UTC window only. • UNet𝜆=2 , yearly, val 2023: model trained on a yearly split (2019–2022), evaluated on a representative sample of 2023 spanning all months and seasons. • UNet𝜆=2 , weekly, on 2023: weekly-trained model evaluated exclusively on the year 2023. • UNet𝜆=2 , yearly, on Nov–Dec: yearly-trained model evaluated on November–December only. • UNet𝜆=2 , yearly, val 2023, {season}: yearly-trained model evaluated on 2023, stratified by season.
31 Table S4: Additional downscaling validation experiments: aggregate metrics for the pipeline using the bias correction model (UNet, 𝜆time = 2) under different train–test splits and evaluation conditions. All metrics are averaged over lead times 6 h to 6 days. Model
crps
mse
mse wet
csi
avfss 5
eS
eA
eL
UNet𝜆=2 UNet𝜆=2 , weekly, new Hours UNet𝜆=2 , weekly new Hours, only 18 UNet𝜆=2 , yearly, on Nov-dec UNet𝜆=2 , yearly, val 2023 UNet𝜆=2 , weekly, on 2023 UNet𝜆=2 , yearly, val 2023, winter UNet𝜆=2 , yearly, val 2023, spring UNet𝜆=2 , yearly, val 2023, summer UNet𝜆=2 , yearly, val 2023, autumn
0.434 0.374 0.500 0.352 0.507 0.533 0.305 0.422 0.754 0.493
3.570 3.011 4.836 1.642 3.833 4.937 1.165 2.065 8.338 3.164
10.328 10.039 13.733 3.802 12.968 12.226 2.380 7.924 31.521 6.456
0.322 0.306 0.325 0.376 0.361 0.303 0.392 0.409 0.302 0.343
0.449 0.423 0.443 0.505 0.471 0.423 0.546 0.544 0.411 0.450
0.473 0.596 0.602 0.137 0.541 0.062 0.398 0.499 0.698 0.199
-0.313 -0.276 -0.392 -0.506 -0.292 -0.515 -0.244 -0.445 -0.522 -0.325
0.211 0.205 0.201 0.195 0.235 0.238 0.162 0.158 0.244 0.221
Table S5: MSE on wet pixels (observed precipitation > 0.1 mm) resolved by lead time for the additional validation experiments. The seasonal decomposition of the yearly model on 2023 (bottom four rows) reveals a ×13 variation between winter and summer, illustrating the dominant role of convective regimes in driving prediction difficulty. Model UNet𝜆=2 UNet𝜆=2 , weekly, new Hours UNet𝜆=2 , weekly new Hours, only 18 UNet𝜆=2 , yearly, on Nov-dec UNet𝜆=2 , yearly, val 2023 UNet𝜆=2 , weekly, on 2023 UNet𝜆=2 , yearly, val 2023, winter UNet𝜆=2 , yearly, val 2023, spring UNet𝜆=2 , yearly, val 2023, summer UNet𝜆=2 , yearly, val 2023, autumn
6h
12h
18h
1d
2d
3d
4d
5d
6d
Avg
9.224 9.228 12.768 2.797 12.314 10.686 1.582 7.209 31.539 5.253
9.437 10.225 14.704 2.935 12.170 10.666 1.859 6.888 31.101 5.251
9.520 9.620 13.348 3.129 12.406 10.961 1.746 7.124 31.339 5.767
9.578 9.730 13.591 3.290 12.438 11.483 2.031 7.012 30.981 6.152
10.040 9.679 13.168 3.518 12.698 11.887 2.434 7.897 31.076 5.868
10.766 10.118 13.331 3.664 12.885 13.308 2.310 7.846 31.440 6.356
11.087 10.395 14.023 4.411 13.278 13.390 2.811 8.456 31.308 7.020
11.408 10.640 14.004 5.074 13.913 13.746 3.123 9.115 32.419 7.384
11.889 10.718 14.660 5.399 14.607 13.910 3.522 9.765 32.488 9.049
10.328 10.039 13.733 3.802 12.968 12.226 2.380 7.924 31.521 6.456
1) Temporal autocorrelation of precipitation fields To quantify the temporal dependence between consecutive samples, we compute the temporal autocorrelation function (ACF) of the 6-hour accumulated precipitation fields from the CombiPrecip observational dataset over Switzerland, at the 12-hour sampling cadence corresponding to the actual spacing between consecutive samples in the dataset. We consider two complementary ACF estimates. The spatial-mean ACF is computed on the domain-averaged precipitation time series and captures the persistence of large-scale synoptic patterns. The field-level ACF is computed independently at each grid point and then averaged spatially: it captures the local decorrelation relevant to the pixel-level predictions made by the model. Results are shown in Figure S6. The spatial-mean ACF exhibits a lag-1 (12 h) autocorrelation of 0.42 and falls below the 0.2 threshold at approximately 36 hours (1.5 days), reflecting the persistence of synoptic-scale weather regimes over the domain. In contrast, the field-level ACF decorrelates much faster, dropping below 0.2 within 12 hours (0.5 days). This indicates that while the domain-averaged precipitation signal may persist for 1–2 days, the local pixel-level structure— which is what the bias correction and downscaling models must predict—changes substantially between consecutive samples. At the weekly split boundary, the minimum gap between the end of the last training accumulation window and the start of the first test accumulation window is 6 hours. The gap between consecutive sample centers is 12 hours. Since the
32 Table S6: CRPS resolved by lead time for the additional validation experiments. The weekly-trained model evaluated on 2023 (weekly, on 2023) and the yearly-trained model (yearly, val 2023) track each other closely across all lead times, confirming that the weekly split does not produce inflated short-range scores through temporal autocorrelation. Model
6h
12h
18h
1d
2d
3d
4d
5d
6d
Avg
UNet𝜆=2 UNet𝜆=2 , weekly, new Hours UNet𝜆=2 , weekly new Hours, only 18 UNet𝜆=2 , yearly, on Nov-dec UNet𝜆=2 , yearly, val 2023 UNet𝜆=2 , weekly, on 2023 UNet𝜆=2 , yearly, val 2023, winter UNet𝜆=2 , yearly, val 2023, spring UNet𝜆=2 , yearly, val 2023, summer UNet𝜆=2 , yearly, val 2023, autumn
0.373 0.330 0.449 0.293 0.435 0.460 0.242 0.368 0.681 0.397
0.383 0.342 0.472 0.303 0.450 0.465 0.259 0.369 0.702 0.420
0.388 0.332 0.454 0.311 0.461 0.485 0.257 0.384 0.706 0.444
0.392 0.337 0.461 0.321 0.475 0.507 0.276 0.397 0.718 0.458
0.417 0.356 0.485 0.327 0.486 0.513 0.290 0.412 0.737 0.453
0.450 0.374 0.492 0.343 0.511 0.576 0.287 0.439 0.774 0.488
0.479 0.406 0.540 0.384 0.545 0.581 0.329 0.454 0.803 0.539
0.494 0.438 0.568 0.424 0.559 0.599 0.375 0.457 0.805 0.549
0.532 0.452 0.581 0.462 0.635 0.607 0.425 0.517 0.857 0.687
0.434 0.374 0.500 0.352 0.507 0.533 0.305 0.422 0.754 0.493
field-level ACF is already below 0.2 at a 12-hour lag, the effective independence between train and test samples at the boundary is high at the spatial scale relevant to the model.
Fig. S6: Temporal autocorrelation function (ACF) of 6-hour accumulated precipitation from CombiPrecip over Switzerland. Left: Full ACF up to 15 days. Right: Zoom on the first 48 hours. The spatial-mean ACF (blue) captures synoptic-scale persistence and decorrelates at approximately 36 h. The field-level ACF (green), computed per pixel and then spatially averaged, decorrelates within 12 h. Vertical lines indicate the minimum train–test boundary gap (6 h, dotted) and the sample-center gap (12 h, dash-dotted). The red dashed line marks the ACF = 0.2 threshold.
2) Neutral-terrain comparison: weekly vs. yearly split To empirically verify that the weekly split does not inflate model performance through temporal leakage, we compare two models on a common, previously unseen test period. Both models use the same architecture (UNet, 𝜆time = 2) and hyperparameters; only the splitting strategy differs. The key comparison involves three rows in Table S4: 1. UNet𝜆=2 , weekly: the reference model on its own weekly validation set (CRPS = 0.434, MSEwet = 10.3). 2. UNet𝜆=2 , weekly, on 2023: the same weekly-trained model, evaluated exclusively on 2023 (CRPS = 0.533, MSEwet = 12.2).
33 3. UNet𝜆=2 , yearly, val 2023: the yearly-trained model, evaluated on the same 2023 period (CRPS = 0.507, MSEwet = 13.0). If the weekly split had produced an artificially strong model through temporal leakage, one would expect it to degrade substantially on the neutral 2023 test set, performing significantly worse than a yearly-trained model that has never relied on neighboring training days. Instead, the two models perform comparably on 2023, with the weekly-trained model even slightly better on MSEwet (12.2 vs. 13.0). This indicates that the weekly split does not produce an inflated model. The degradation from the weekly validation set (CRPS = 0.434) to the 2023 evaluation (CRPS = 0.533) is therefore attributable to a distribution shift rather than to the disappearance of leaked autocorrelation. This interpretation is strongly supported by the seasonal decomposition of the yearly model’s 2023 metrics (Table S5): MSEwet ranges from 2.4 in winter to 31.5 in summer—a ×13 factor—revealing that performance is dominated by the intrinsic difficulty of convective precipitation regimes rather than by the choice of splitting strategy. The result on November–December alone (UNet𝜆=2 , yearly, on Nov–Dec: CRPS = 0.352, MSEwet = 3.8) further illustrates this seasonal effect: evaluating on a period dominated by stratiform precipitation yields substantially better metrics, regardless of the splitting strategy used. Lead-time-resolved results (Tables S5 and S6) confirm that the degradation pattern is consistent across all lead times: the weekly-on-2023 and yearly-on-2023 models track each other closely from 6 h through 6 days, with both exhibiting the expected monotonic increase in error with lead time. 3) Intrinsic difficulty across accumulation windows To disentangle the effect of the diurnal cycle from potential generalization gaps, we measure the intrinsic prediction difficulty of each 6-hour accumulation window by comparing raw AIFS forecasts directly against CombiPrecip observations coarsened to the AIFS grid via bilinear interpolation. This comparison involves no trained ML model and thus provides a model-independent baseline of difficulty. Table S7 reports the MSE on wet pixels (observed precipitation > 0.1 mm) for each window, averaged over all lead times. Table S7: Intrinsic difficulty of raw AIFS predictions vs. observations by accumulation window. MSEwet is computed on the AIFS grid after coarsening observations via bilinear interpolation. The difficulty ratio is relative to the easiest window (00–06 UTC). Window (UTC) 00–06 06–12 12–18 18–00
Role
MSEwet
Difficulty ratio
held-out training held-out training
13.6 15.9 26.6 25.1
×1.00 ×1.17 ×1.96 ×1.85
The 12–18 UTC window (afternoon convection) is nearly twice as difficult as the 00–06 UTC window (nighttime, predominantly stratiform precipitation). Crucially, the 18–00 UTC window—which is included in the training set— exhibits a comparable level of difficulty (×1.85), reflecting the persistence of convective activity into the evening hours. The training set therefore already exposes the model to the challenging convective regime. This difficulty structure explains the results observed when evaluating the weekly-trained model on the held-out time windows (Table S4): • UNet𝜆=2 , weekly, new Hours (both held-out windows combined): CRPS = 0.374, MSEwet = 10.0. This is better than the reference model’s own validation score (CRPS = 0.434), because the easy 00–06 window dominates the average. • UNet𝜆=2 , weekly new Hours, only 18 (12–18 UTC window alone): CRPS = 0.500, MSEwet = 13.7. The degradation relative to the reference is consistent with the ×1.96 intrinsic difficulty ratio of this window compared to the easiest window. The model thus does not exhibit a generalization failure on the held-out windows; rather, its performance reflects the diurnal cycle of precipitation predictability over Switzerland.
34 S5. Additional Visualizations: November 2023 Heavy Precipitation Event We present test diffusion samples from the yearly trained SwAIther-Precip pipeline based on the UNet𝜆=2 bias correction model, for a particularly heavy precipitation event from November 12th to November 15th , 2023. For visualization purposes, we show the first ensemble sample (member 0) out of 12 diffusion samples. Each figure shows the downscaling output across all forecast lead times.
Fig. S7: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (1/7).
Fig. S8: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (2/7).
Fig. S9: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (3/7).
35
Fig. S10: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (4/7).
Fig. S11: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (5/7).
Fig. S12: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (6/7).
S6. Forecasting verification metrics: detailed definitions We evaluate the performance of our bias correction models as well the resulting downscaling models using a comprehensive set of metrics that capture different aspects of precipitation forecast quality. These include point-wise error metrics, categorical event-based metrics, and spatial focused metrics. 1) Pointwise and categorical verification metrics Point-wise metrics measure the magnitude of prediction errors at each grid point independently. Since precipitation distributions are typically dominated by dry grid points, we compute an MSEwet restricted to wet pixels, meaning pixels where the precipitation is larger than 0.1 mm/6h.
36
Fig. S13: Downscaling samples by lead time (member 0) for the November 2023 heavy precipitation event (7/7).
Table S8: Interpretation of FSS at different spatial scales Neighborhood 𝑛=1 𝑛=5 𝑛 = 17
Scale
Interpretation
Pixel-level ∼25 km ∼85 km
Exact location skill (strict) Small displacement tolerance Synoptic-scale skill
Event-based metrics evaluate the model’s ability to correctly detect precipitation occurrence by treating the forecast as a binary classification problem. Wet pixels are defined by a threshold of 0.1 mm/6h, as mentioned in Section 3d of the main text. This yields a contingency table with four categories: true positives (TP, correctly predicted rain), false positives (FP, predicted rain that did not occur), false negatives (FN, missed rain events), and true negatives (TN, correctly predicted dry conditions). The Critical Success Index (CSI) is then computed to measure overall detection skills: CSI = TP/(TP + FP + FN), ranging from 0 to 1. By excluding true negatives, CSI is well suited for precipitation verification where dry conditions dominate. 2) Spatial Verification Metrics Traditional point-wise metrics suffer from the “double penalty” problem, where spatially displaced but otherwise correct forecasts are penalized twice. Spatial verification metrics address this by evaluating quality within a neighborhood context. Fractions Skill Score (FSS) and member-averaged FSS (avFSS). The FSS (Roberts and Lean 2008) evaluates spatial agreement between forecast and observed precipitation exceedance fields smoothed over a neighborhood of size 𝑛 × 𝑛. We adopt the reformulation of Necker et al. (2024) based on Binary Probabilities (BP): BP𝑖 =
1 ∑︁ 1[𝐹 𝑗 ≥ 𝜏] 𝑛2 𝑗 ∈ N
(S1)
𝑖
The FSS is defined as FSS = 1 − FBS/FBSref , where FBS is the mean squared difference between forecast and observed BP fields, and FBSref is the worst-case reference assuming zero spatial overlap. FSS ranges from 0 to 1, with 1 indicating perfect spatial agreement. For Í 𝑀ensemble forecasts, we employ the member-averaged FSS (avFSS) (Necker et al. 2024), computed as avFSS = 𝑀 −1 𝑚=1 FSS(𝐹𝑚 , 𝑂), which preserves member sharpness unlike FSS of the ensemble mean. We compute FSS and avFSS at three neighborhood sizes: 𝑛 = 1 (pixel-level), 𝑛 = 5 (∼25 km), and 𝑛 = 17 (∼85 km), as summarized in Table S8. Structure-Amplitude-Location (SAL) and Ensemble SAL (eSAL). The SAL diagnostic (Wernli et al. 2008) decomposes forecast error into three independent components. Structure (S) compares the shape and peakedness of precipitation objects (𝑆 = (𝑉𝐹 −𝑉𝑂 )/[0.5(𝑉𝐹 +𝑉𝑂 )], where 𝑉 = 𝑅/𝑅max ); negative values indicate overly peaked fields, positive values overly diffuse ones. Amplitude (A) compares domain-average precipitation (𝐴 = (𝑅𝐹 − 𝑅𝑂 )/[0.5(𝑅𝐹 + 𝑅𝑂 )]);
37 negative indicates underestimation, positive overestimation. Location (L) quantifies spatial displacement based on precipitation-weighted centers of mass. All three components equal zero for a perfect forecast and range within [−2, 2] for S and A, and [0, 2] for L. For ensemble forecasts, we report the ensemble SAL (eSAL) (Radanovics et al. 2018), defined as the median of each SAL component computed across ensemble members: eS = median({𝑆 𝑚 }), eA = median({𝐴𝑚 }), eL = median({𝐿 𝑚 }). 3) Probabilistic Verification Metrics The second stage of SwAIther-Precip employs a diffusion model to generate an ensemble of high-resolution precipitation fields, providing a probabilistic estimate of the downscaled precipitation. Evaluating probabilistic forecasts requires metrics that assess both the reliability (statistical consistency) and sharpness (concentration) of the predicted distributions. We employ the CRPS to to evaluate the quality of our probabilistic predictions. Continuous Ranked Probability Score (CRPS). The CRPS (Matheson and Winkler 1976; Hersbach 2000) is a strictly proper scoring rule that generalizes MAE to probabilistic forecasts. For an ensemble of 𝑀 members {𝑥1 , . . . , 𝑥 𝑀 } and observation 𝑦, it is computed as (Gneiting and Raftery 2007): CRPS(𝐹, 𝑦) =
𝑀 𝑀 𝑀 1 ∑︁ 1 ∑︁ ∑︁ |𝑥 𝑚 − 𝑥 𝑚′ | |𝑥 𝑚 − 𝑦| − 𝑀 𝑚=1 2𝑀 2 𝑚=1 𝑚′ =1
(S2)
where the first term measures ensemble–observation distance and the second rewards spread. CRPS has the same unit as the predicted variable and reduces to MAE for deterministic forecasts. We report the spatial average over all grid points. Probability Integral Transform (PIT). The PIT (Gneiting and Raftery 2007; Dawid 1984) assesses ensemble calibration by evaluating the quantile at which the observation falls within the predicted distribution. For a calibrated forecast, PIT values are uniformly distributed on [0, 1]; deviations reveal underdispersion (U-shaped histogram), overdispersion (dome-shaped), or systematic bias (skewed). Because precipitation has a point mass at zero, we use the randomized PIT (Gneiting and Raftery 2007): PIT ∼ U [𝐹 (𝑦 − ), 𝐹 (𝑦)] (S3) where 𝐹 (𝑦) and 𝐹 (𝑦 − ) are the empirical CDF evaluated at and just below 𝑦, respectively. We aggregate PIT values into histograms with 𝐾 = 20 bins and quantify departure from uniformity using the Kullback–Leibler divergence (Leinonen et al. 2023): 𝐾 ∑︁ ℎ𝑘 ℎ 𝑘 log 𝐷 KL = (S4) 1/𝐾 𝑘=1 where 𝐷 KL = 0 corresponds to perfect calibration. We additionally compute PIT histograms stratified by precipitation intensity (dry, light, heavy) to diagnose regime-dependent miscalibration. 4) Spectral Evaluation Point-wise and categorical metrics do not characterize the spatial scales at which the downscaling pipeline produces reliable precipitation structures. A model may achieve low MSE while systematically under-representing variability at certain scales—a deficiency consequential for hydrological applications. Spectral analysis addresses this by decomposing forecast fields into contributions at each spatial wavelength. Radially-averaged power spectral density. Following Sinclair and Pegram (2005); Pulkkinen et al. (2019), we compute the radially-averaged power spectral density (RAPSD) of log-transformed precipitation fields (𝑃˜ = log(𝑃 + 𝜀), 𝜀 = 0.1 mm/6h). After removing the spatial mean and applying a Tukey tapering window (𝛼 = 0.1), we compute the 2D PSD and radially average it to obtain a 1D spectrum as a function of wavelength 𝜆 = 1/𝑘 (km). Co-masked spectral comparison. Since predicted and observed fields generally have different wet-area extents, a direct spectral comparison would conflate coverage differences with structural differences. We therefore adopt a co-masking strategy: for each time step and ensemble member, both fields are restricted to the intersection of their wet areas (> 0.1 mm/6h) before computing spectra. Time steps where the common wet area covers less than 3% of the domain are excluded. Scale-dependent spectral ratio and effective resolution. The primary diagnostic is the spectral ratio PSDpred (𝜆)/PSDobs (𝜆): a ratio of 1 indicates perfect spectral fidelity, values below 1 indicate under-dispersion, and
38 Table S9: Wavelength bands for scale-dependent spectral analysis Band Large scale Mesoscale Small scale
Wavelength range 100–600 km 20–100 km 2–20 km
Interpretation Synoptic precipitation Frontal bands Convective cells, fine-scale
values above 1 indicate over-dispersion. We compute band-averaged ratios over three wavelength bands, as reported in Table S9. Following Klaver et al. (2020), we define the effective resolution as the wavelength at which the spectral ratio first drops below 0.5, scanning from large to small scales. This provides a single-number characterization of the pipeline’s resolving capability; for reference, NWP models typically achieve effective resolutions of 4–8 times their grid spacing (Skamarock 2004; Klaver et al. 2020). In the context of our pipeline, band-averaged ratios at large and mesoscales primarily reflect the bias correction quality (Step 1), while the small-scale band characterizes the diffusion model’s ability to generate realistic fine-scale texture (Step 2).