ConceptioArchivearXiv CS
arXiv CSopen access

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

arXiv:2607.16050v1 [cs.LG] 17 Jul 2026

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings Yuya Kawakami

Daniel Cayan

Dongyu Liu

University of California, Davis Davis, California, USA

Scripps Institution of Oceanography University of California, San Diego La Jolla, California, USA

University of California, Davis Davis, California, USA

Kwan-Liu Ma

Tom Corringham

University of California, Davis Davis, California, USA

Scripps Institution of Oceanography University of California, San Diego La Jolla, California, USA

Abstract

Keywords

Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to coarse resolution, regional domains, or computationally intensive process-based models unsuitable for daily continental-scale use. We present DELUGE, a multimodal deep learning framework for daily pluvial flood damage prediction at ∼1 km resolution and national scale, trained on spatially and temporally corrected NFIP claims (2017–2022) and structured around the hazard, exposure, and vulnerability components of disaster risk. Rather than blanket coverage of the Conterminous United States (CONUS), we model the top 100 highest-claim 75 km cells, distributed nationwide and accounting for ∼81% of total pluvial flood claims. Our architectural novelty is a pair of parametric modules in the hydrometeorology branch, a Value Modulator and a Temporal Modulator, conditioned on terrain descriptors and AlphaEarth foundation-model embeddings, that expose directly inspectable hydrological response parameters and provide architecture-level interpretability-by-design. Under a spatial block holdout, DELUGE outperforms tuned Random Forest, XGBoost, and LightGBM baselines by 9% to 30% on a dollar-weighted area under the precisionrecall curve (PR-AUC), a metric that emphasizes the rare, high-cost claims of greatest operational interest. Beyond DELUGE, we argue this interpretable conditioning scheme is a transferable pattern for integrating foundation-model embeddings into other geospatial prediction tasks.

GeoAI, Flood Damage Prediction, Pluvial Flooding, Interpretability, Geospatial Foundation Models

CCS Concepts • Computing methodologies → Model development and analysis; • Applied computing → Earth and atmospheric sciences. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. Conference acronym ’XX, Woodstock, NY © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/XXXXXXX.XXXXXXX

ACM Reference Format: Yuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu Ma, and Tom Corringham. 2018. DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’XX). ACM, New York, NY, USA, 18 pages. https://doi.org/XXXXXXX.XXXXXXX

1

Introduction

Flooding is among the costliest natural hazards in the conterminous United States (CONUS) [23], with damage dynamics that vary sharply by topography, drainage infrastructure, and climate regime [14]. Among flood types, pluvial flooding (inundation driven by intense rainfall exceeding local drainage capacity) is a particularly difficult prediction target [51, 58]. Anticipating pluvial damage at a resolution operationally meaningful to insurers, emergency managers, and municipal planners is both an open scientific problem and is of immediate societal value. Despite this urgency, pluvial flood damage prediction lags behind its riverine and coastal counterparts [58]. Riverine flooding benefits from a dense network of stream gauges that provide a relatively clean, continuous training signal [11, 50], and coastal flooding is similarly well-tracked by tide and surge gauges [7, 76]. Pluvial flooding, on the other hand, has a noisy observational backbone, and its damage signal lives in scattered insurance claims [24] and citizen reports [47, 54], each carrying nontrivial spatial and temporal uncertainty. CONUS-scale pluvial products do exist, but they occupy a different operating regime than ours. FEMA flood maps [22] are static delineations of long-term flood hazard zones rather than dynamic estimates of damage, while CONUS-scale hydrodynamic models [7, 73] simulate physical inundation at high fidelity but are computationally heavy and produced for fixed design scenarios rather than on a daily basis. Neither delivers a lightweight, daily, ∼1 km prediction of insured pluvial damage, the regime we target. We close this gap with DELUGE (Daily Estimation of pLU vial flood damaGE), to our knowledge the first end-to-end learned system to predict daily, ∼1 km insured pluvial flood damage across the

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

highest-claim regions of CONUS through interpretable conditioning on Earth-observation foundation-model embeddings. DELUGE is a multimodal CNN whose local inductive bias matches the locally driven nature of pluvial dynamics, in contrast to the catchmentscale connectivity behind GNN-based riverine approaches [35, 61]. We train DELUGE on NFIP claims as a damage proxy, after applying temporal and spatial uncertainty corrections that refine each claim’s timing and footprint. Under a spatial block holdout protocol that partitions 75 km grid cells into a train/test split, DELUGE achieves a PR-AUC of ∼0.24, well above the ∼0.0025 no-skill baseline implied by the task’s ∼0.25% positive rate, and outperforms tuned Random Forest, XGBoost, and LightGBM baselines commonly used in claims-based flood-damage prediction, with the largest gains on high-cost claims. Beyond its predictive performance, DELUGE’s architectural novelty lies in two parametric modules within the hydrometeorology branch. The Value Modulator learns a per-sample monotone warp of each hydrometeorological input, and the Temporal Modulator learns a per-sample temporal-response kernel that causally convolves the hydrometeorological time series. Both modules are conditioned on local place characteristics, encoded through terrain descriptors and Google AlphaEarth foundation-model embeddings [13], which supply a dense, learned representation of the built and natural environment that varies continuously across CONUS. Rather than concatenating these embeddings into a black-box encoder, we route them through the physics-shaped modulators, whose parameters carry direct hydrological meaning such as response peak lag in hours and decay timescale. This makes the modulators a principled and interpretable way to consume an otherwise opaque foundation embedding, yielding interpretability-bydesign [60] where ConvLSTM or Transformer-style encoders would require post-hoc attribution approximations like SHAP [45]. Thus our interpretability is at the architecture level, exposing the perlocation hydrological response DELUGE has learned, rather than a per-prediction explanation. Our main contributions are: (1) DELUGE, the first end-to-end learned system for daily pluvial flood damage prediction at ∼1 km across the highest-claim regions of CONUS, outperforming tree-based baselines standard in the pluvial flood prediction literature, with the largest gains on high-cost claims. (2) Two uncertainty-correction procedures for NFIP claims, a precipitation guided correction of claim dates and a spatial refinement that intersects redacted geographies to localize damage to higher resolution. (3) The Value Modulator and Temporal Modulator, terrain- and AlphaEarth-conditioned parametric modules whose learned per-patch parameters are themselves hydrologically meaningful, yielding interpretability-by-design of Geospatial Foundation Model integration. (4) Inspection of the learned modulator parameters via geographyblind clustering, showing they recover hydrologically coherent regimes consistent with physical expectation.

2

Related Works

Flood Hazard Modeling and Process-Based Approaches. Flood risk mapping has traditionally relied on process-based hydrological and hydraulic models that simulate fluid physics over high-resolution

Kawakami et al.

topography [5, 37]. These models underpin FEMA flood maps and CONUS-scale ambient-flood products [7, 22, 73, 75], providing 30 mto-street-level inundation footprints valuable for insurance and long-term planning, often at decadal scales. For damage estimation specifically, the insurance industry has developed catastrophe (CAT) models that integrate hazard, exposure, and vulnerability for detailed risk assessment [77]. However, neither of these approaches is suited to daily, dynamic damage prediction. Ambient-flood products characterize long-run risk rather than the daily prediction, and CAT models are computationally expensive to run at scale [6, 35]. Machine Learning for Flood Prediction. Machine learning approaches address the operational gap left by process-based models by learning directly from large archives of hydrological, topographical, and socioeconomic data [2, 17, 41, 52]. Tabular tree-based methods (Random Forest, XGBoost, LightGBM) are the dominant choice in the flood damage prediction literature [2, 17, 26, 76], typically operating on aggregated features at the county or grid-cell level. A parallel line of work applies sequence and graph models (LSTM, GNN) to riverine and flash-flood prediction [35, 48, 50, 61, 66, 67], achieving strong results, but the targets are typically river gauge measurements or hazard indicators rather than insured damage directly. A further body of remote-sensing work maps flood extent from satellite imagery [10, 54, 56], but likewise recovers inundation footprints rather than damage. Among damage-targeted ML works, most are constrained to regional domains (a single county, or city) where local data permit detailed modeling [43, 63, 78], leaving the CONUS-wide daily pluvial damage problem largely unaddressed. Pluvial Flood Damage Prediction from Claims. Pluvial flooding is a particularly difficult prediction target [51, 58], due to the highly localized natures of impacts and lack the dense observational infrastructure of stream and tide gauges that supports riverine and coastal flood predictions. NFIP claims are the most comprehensive source of insured pluvial damage in the United States [24] and have become the de facto damage signal for data-driven CONUS-scale work [26, 76]. However, claims data are known to carry nontrivial spatial and temporal uncertainty. Data redaction to protect personally identifiable information limits localization to Census Tract or ZIP polygons, reported dates frequently misalign with precipitation events due to reporting lags and mislabeling [53, 62, 74], and uninsured losses are systematically absent (∼2/3 of total losses [4]). Past work has largely treated these uncertainties as fixed properties of the dataset. Here, we instead build an explicit temporal and spatial uncertainty correction pipeline (§3.1.1, §3.1.2) to refine each claim before it enters the model, which we view as a prerequisite for using NFIP claims as a daily training signal. GeoAI Foundation Models and Interpretability. Earth-observation foundation models are increasingly becoming backbones of geospatial analysis tasks [1, 65, 79]. Efforts like Google AlphaEarth [13] provide dense, high-resolution learned embeddings that implicitly encode the built and natural environment and spatial context. A growing body of work also probes what these embeddings represent about physical space [8, 9, 55]. However, less attention has been paid to how downstream models should consume them. The default pattern, which is to concatenate embeddings into a blackbox encoder, makes it difficult to verify whether the geospatial prior

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

is being used in a physically meaningful way. This raises the same interpretability concerns that motivate broader calls for explainable GeoAI [30, 59, 64, 79] and interpretable flood models [66], and that Rudin [60] has addressed by advocating inherently interpretable architectures over post-hoc attribution. DELUGE’s interpretable conditioning modules are designed to fill exactly this gap. Rather than treating foundation-model embeddings as opaque feature vectors, we route them through parametric, physics-shaped modules whose learned parameters are themselves hydrologically meaningful and inspectable.

Grid Cell

3 Data 3.1 Prediction Target: FIMA NFIP Claims Our prediction target is derived from the OpenFEMA FIMA NFIP Redacted Claims Dataset v2 [24], which records every NFIP insurance claim filed under the National Flood Insurance Program, including the reported dateOfLoss, redacted location (Census Tract, ZIP code, and nearest 0.1◦ × 0.1◦ grid cell), causeOfDamage category, and paid amounts for building and contents coverage. As we target pluvial claims, we filter the dataset to include only those claims defined as causeOfDamage = Accumulation of rainfall or snowmelt, which serve as the supervisory signal for our prediction task. NFIP data, while the most comprehensive source of insured flood damage in CONUS, is known to have uncertainties and errors in both space and time [53, 62, 74]. To improve learning dynamics and model validation, we apply two correction methods, temporal and spatial, to mitigate these uncertainties. 3.1.1 Temporal Uncertainty Correction. The dateOfLoss field in the NFIP claims data is known to be noisy where reported dates of loss often fail to align with precipitation events in the historical precipitation records. These discrepancies may arise from reporting lags, human error during claim filing, and misclassification of causeOfDamage, where fluvial or coastal damages are mislabeled as pluvial [62]. To mitigate this noise, we apply the following correction to the dateOfLoss field 𝑑: (1) Compute the daily maximum hourly accumulated precipitation from NOAA AORC [21] for day 𝑑 at the location of the claim. (2) If the daily maximum hourly accumulated precipitation on day 𝑑 exceeds 10 mm, accept dateOfLoss as reported. (3) Otherwise, search the window 𝑑 ′ ∈ {𝑑 − 3, . . . , 𝑑 + 3} and select the day with the largest daily maximum hourly precipitation. If this maximum exceeds 10 mm, update dateOfLoss to 𝑑 ′ . (4) If no day’s daily maximum hourly precipitation in the window exceeds 10 mm, discard the claim as erroneous. We adopt 10 mm as a conservative threshold for maximum hourly accumulated precipitation, sitting between two reference points. National Weather Service Flash Flood Guidance (FFG) thresholds are most predictive for most areas when hourly precipitation exceeds ∼25.4 mm [34], while the American Meteorological Society’s Glossary of Meteorology defines heavy precipitation as 7.62 mm in one hour [3]. With this processing, 81.1% of claims are accepted as reported, 12.9% are shifted within ±3 days, and 6.0% are discarded. We place the full table of date shifts in the Appendix. 3.1.2 Spatial Uncertainty Correction. The address-level claims data is not publicly available and the redacted claims data is only identifiable by its Census Tract, Zip Code and the nearest 0.1◦ × 0.1◦ degree

Patch

~8km Context Patc

Center Pixel ~1.1k Prediction: NFIP pluvial damage occurence on day D

~8k Figure 1: Our study region in blue. We select the Top 100 75×75 km grid cells for NFIP pluvial flood damage claims, accounting for ∼81% of claims from 2017–2022. After rasterizing NFIP claims (see §3.1.1,§3.1.2), training samples take the form of (pixel, day) tuples, and DELUGE takes input a ∼8x8km patch centered around the pixel as spatial context. 75km

cell. To mitigate this spatial uncertainty of the location of the claim, for every claim, we retrieve the intersection of the reported Census Tract, Zip Code, and the 0.1◦ × 0.1◦ cell centered on the reported coordinate, yielding a polygon for each claim. After this spatial uncertainty correction, our median polygon area for each claim is ∼ 1.7km2 with a mean area of ∼ 7.7km2 due to a significant right tail of large rural polygons.

3.2

Study Region and Period

We segment CONUS into 75×75 km grid cells and select the top 100 cells by pluvial NFIP damage claim count, yielding a nationally distributed set of the most flood-affected regions rather than blanket coverage of CONUS. These cells span CONUS, while concentrating in the Northeast and Gulf regions of the United States (Fig. 1), and together account for ∼81% of all pluvial NFIP claims from 2017– 2022. Extending to the top 200 cells would raise claim coverage only modestly to ∼91% while roughly doubling the dataset, so we adopt the top 100 as a balance between capturing the large majority of pluvial claims and keeping storage and compute tractable. Within each 75×75 km grid cell, we further discretize space into a 64×64 raster of pixels (∼1.17 km on a side), which constitute the unit of spatial prediction throughout the remainder of the paper. We restrict our study period to 2017–2022, the longest contiguous interval over which all data sources have full-year coverage.

3.3

Dataset Construction

The unit of binary prediction is a (pixel, day) tuple over our study region and period, i.e., was there a pluvial NFIP damage claim reported at pixel 𝑝 on day 𝑑? First to collect our positive class, we rasterize each claim’s spatial-uncertainty polygon (§3.1.2) onto the 64×64 grid, labeling every intersecting pixel as positive. To prevent

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

claims with large polygons from dominating the loss, we associate each claim with a confidence score: 𝑐𝑐𝑙𝑎𝑖𝑚 = min(1, 𝐴pix /𝐴claim ) with 𝐴pix = 1.172 km2 . To account for pixels with many intersecting claims, we calculate and define our confidence at each pixel 𝑐 𝑝𝑖𝑥𝑒𝑙 as the complement of the product of each claim’s confidence Î score: 𝑐 𝑝𝑖𝑥𝑒𝑙 = 1 − 𝑖 ∈claim (1 − 𝑐𝑖 ). This operation captures the intuition that multiple claims intersecting a pixel should increase the confidence of a claim for that pixel. For the negative class, exhaustively enumerating every pixel-day in our study period would yield a label distribution dominated by trivially negative dry days, further exacerbating class imbalance. We therefore construct the dataset along three data types drawn from two contrastive axes: positives, spatial negatives (same-day, different-pixel), and temporal negatives (different-day, same-grid). Positives: Pixel-days containing at least one NFIP pluvial claim after the temporal and spatial uncertainty corrections described above, along with a confidence score 𝑐. Spatial Negatives: Pixel-days without a claim on grid-days that contain at least one claim somewhere within the same 75×75 km grid cell. Temporal Negatives: Pixeldays on grid-days where the spatial mean of the 1-hour maximum accumulated precipitation from NOAA AORC exceeds 10 mm and no NFIP claim was filed anywhere in the grid cell. We reuse the 10 mm threshold from our temporal uncertainty correction (§3.1.1) so that meaningful precipitation events are defined consistently across the pipeline. Despite this undersampling of temporal negatives, the dataset exhibits significant class imbalance with a positive rate of ∼0.25% across the study region and period.

4

DELUGE Architecture

Pluvial flooding is often a localized event with dynamics that depend strongly on local conditions such as topography, land use, and drainage infrastructure. Although pluvial events, especially streetlevel events, may occur at sub-kilometer scales [46] beyond the resolution of our modeling, we provide the model with a larger ∼8 × 8 km spatial context centered on each sample (a 7 × 7 pixel patch) for two reasons. First, prior work has identified the spatial scale at which pluvial flooding develops [19, 72]. The effective extent varies across locations, but our ∼8 × 8 km context covers the scale of pluvial flooding identified in the literature. Second, pluvial flood events are often driven by convective storms [40] that deposit large amounts of precipitation in a localized area. These extreme precipitation events are less reliably captured in the NOAA AORC archive [21], and we expect their reported locations to carry variable spatial noise even with the latest observational and radar-based products. A larger spatial context lets the model associate observed damage with the precipitation pattern in the surrounding area, which is more reliably recorded than the storm’s exact center. These considerations motivate DELUGE, the multimodal CNN-based model described in the following sections.

4.1

Overview

We predict per-pixel binary pluvial flood insurance claims occurrence at ∼1.17 km resolution across the highest-claim regions of CONUS (§3.2) by fusing four modality branches (hydrometeorology, terrain, exposure, and vulnerability) into a per-pixel probability

Kawakami et al.

(Figure 2). With respect to the standard triad in disaster risk modeling, the hydrometeorology and terrain branches capture the hazard, while the exposure and vulnerability branches capture exposure and vulnerability respectively. Our target is the center pixel of a 7 × 7 patch (∼8 × 8km) and all branches operate on this patch. 4.1.1 Supporting interpretability. Predictive accuracy alone is insufficient for models intended to inform real-world flood-risk decisions. Insurers and emergency managers must be able to inspect and justify a model’s reasoning before acting on its outputs, especially under the high financial stakes of flood damage. Prior work in deep-learning based flood and hydrological modeling has emphasized this same need for interpretability [38, 39, 66]. We target this need at the architecture level, designing modulators (§4.3 and 4.4) whose hydrological response parameters are directly inspectable, a distinct goal from per-prediction interpretability. We expose this interpretability via the hydrometeorology branch, the most important modality for prediction, where we develop two terrain and foundation model-conditioned modules that modulate the hydrometeorology signal along its two axes. First, the Value modulator (Section 4.3) learns a per-patch, per-channel monotonic warp of the hydrometeorology value range (value axis) and then the Temporal modulator (Section 4.4), which learns a per-patch mixed gamma-kernel hydrological response over the hydrometeorology lookback window (time axis). Both modules are terrain and foundation model-conditioned, so their parameters vary spatially as explicit model outputs (Fig. 3). Conditioned on the local place characteristics, they answer two questions: which range of values in the hydrometeorology signal is most relevant to flood damage (Value Modulator), and what temporal pattern of the hydrometeorology signal is most relevant to flood damage (Temporal Modulator). A similar module to our Value Modulator was offered in [66] via Kolmogorov-Arnold Networks (KAN), but here we allow the model to learn location-specific warps that are conditioned on each local terrain and foundation model features. This is a deliberate interpretability-by-design choice rather than the more common path of pairing a flexible black-box encoder (ConvLSTM, Transformer) with post-hoc attribution (SHAP, integrated gradients, attention visualization) [60], allowing us to inspect the model’s reasoning directly from hydrologically meaningful, spatially explicit modulator parameters instead of generic feature importance scores.

4.2

Data Inputs

Predicting flood damages requires accounting for the three components of risk, hazard, exposure, and vulnerability, and we gather a comprehensive set of inputs to capture each. Hydrometeorological hazard is drawn from hourly NOAA AORC accumulated precipitation [21] and 3-hourly NOAA National Water Model Retrospective volumetric soil moisture and snow water equivalent [18]. Terrain hazard combines the Topographic Wetness Index [29], Height Above Nearest Drainage [20], Global Curve Number under three antecedent runoff conditions [33], NLCD fractional impervious surface [70], and POLARIS soil properties (saturated hydraulic conductivity, clay percentage, and saturated water content) [15]. Exposure and vulnerability come from FIMA NFIP active policy counts and coverage [25] and per-structure attributes from the

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

Data Sources Pluvial NFIP Claims HydrometerologyQ NOAA AORC & NWM AlphaEarth (AEF) HAND, TWI,Q POLARIS, GCN NFIP Active Policies

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Modality Encoders

Dataset Constructio a) Spatial Uncertainty Correction

c) Data Sampling

Terrain & AlphaEarth (AEF) as Modulators

Damage Days Raw NFI Claims

Triangulat& to polygon

Rasterize toQ ~1km grid

b) Temporal Uncertainty Correction

d-3

National StructureQ Inventory

... Reported1 ... d+3 day d Corrected to d-1

Hydromet-Hazard Encoder

1hr Accumulated Precipitation, Soil Moisture, Snow Water Equivalent SpatialQ Positives Negatives

Terrain-Hazard Encoder

Non Damage Days

Exposure Encoder

HAND, TWI, GCN, AEF, POLARIS

Cross-Attention Fusion a) Project to shared dimension

b) Add modality embedding c) Multi-head self-attention across modality tokens

# Active NFIP Policies, # NSI Structures

Vulnerability Encoder

NSI: % Structs. in SFHA Zone, % Structs. Mobile home, % Structs. with basement, % Structs. with Slab Foundation

TemporalQ Negatives

Classifier Head 2-Layer MLPÁ

P(flood damage)

Figure 2: DELUGE architecture. DELUGE is a multimodal CNN-based model that integrates various data modalities, including meteorological, terrain-hazard, vulnerability, and exposure features, to predict pluvial flood damage at a daily and ∼1km resolution across the highest-claim regions of CONUS. The model consists of separate CNN branches for each modality, which are then fused together to produce the final damage prediction. The hydrometeorology branch is described in detail in Fig. 3. USACE National Structure Inventory [69] (structure count, foundation types, mobile-home share, Special Flood Hazard Area [SFHA] membership, first-floor elevation statistics). Finally, to complement these and to condition our modulators, we include 64-dimensional Google AlphaEarth Foundation (AEF) embeddings at 10 m resolution [13], which implicitly encode land cover and the built environment. Per-source descriptions, exact variables used, and feature normalization details are deferred to the Appendix. For each sample, a pixel-day tuple, we extract data for our modalities from the 7 × 7 pixel patch centered on the prediction pixel, bilinearly interpolating to match our ∼1.17 km resolution regular grid. For inputs that are natively offered at a higher resolution, all sources except for the hydrometeorology and NFIP active policies, we extract the data at 56 × 56 resolution. For our hydrometeorology inputs, we extract a 𝑇 = 73-hour input window spanning Day 𝐷 − 2 through Day 𝐷 (inclusive). The temporal modulator (§4.4) convolves this window with an 𝐿 = 48-hour kernel, and the temporal reduction selects a peak-response hour within Day 𝐷. The two are sized so that every reduction hour in Day 𝐷 has its full 𝐿 = 48-hour kernel reach contained inside the 𝑇 = 73-hour input, requiring no padding. We choose a 48-hour time window by considering the typical temporal dynamics of pluvial flooding, which can occur rapidly within hours of heavy rainfall, but also considering the need to capture antecedent conditions like soil moisture that can influence flood risk [12, 34, 68]. A 48-hour kernel allows us to capture both the immediate impact of precipitation and the influence of prior hydrometeorological conditions, providing a comprehensive view of the factors contributing to pluvial flooding.

4.3

Value Modulator

Raw hydrometeorology values often span large ranges, and their flood-relevance is nonlinearly distributed across that range. For instance, a soil moisture measurement of 0.2 m3 m−3 may be more relevant to pluvial flood dynamics in one location than in another. Standard ML pipelines typically apply a global transformation (e.g. min-max standardization or a log transform) to the raw values, but we expect the value-relevance relationship to vary across locations, so such a one-size-fits-all transformation may miss this local variability. This motivates our Value Modulator, which learns a per-patch, per-channel transformation of the raw climate values,

conditioned on the local terrain and foundation model features which we call a warp. Warp definition. For each hydrometeorology channel 𝑣 and patch 𝐾 𝑖, the value modulator produces 𝐾 = 16 width parameters {𝑢 𝑣,𝑖,𝑘 }𝑘=1 𝐾 and 𝐾 = 16 height parameters {ℎ 𝑣,𝑖,𝑘 }𝑘=1 , each softmax-normalized Í Í so 𝑘 𝑢 𝑣,𝑖,𝑘 = 𝑘 ℎ 𝑣,𝑖,𝑘 = 1. We choose 𝐾 = 16 as a balance between expressivity and simplicity, finding 16 control points sufficient to capture the value-relevance relationship. These control points define cumulative breakpoints on the input and output axes, so that Í Í 𝑋 𝑣,𝑖,𝑘 = 𝑘𝑗=1 𝑢 𝑣,𝑖,𝑗 and 𝑌𝑣,𝑖,𝑘 = 𝑘𝑗=1 ℎ 𝑣,𝑖,𝑗 ,with 𝑋 𝑣,𝑖,0 = 𝑌𝑣,𝑖,0 = 0 and 𝑋 𝑣,𝑖,𝐾 = 𝑌𝑣,𝑖,𝐾 = 1. The warp 𝜙 𝑣,𝑖 : [0, 1] → [0, 1] is then a monotonic piecewise-linear map satisfying 𝜙 𝑣,𝑖 (𝑋 𝑣,𝑖,𝑘 ) = 𝑌𝑣,𝑖,𝑘 , with slope 𝑚 𝑣,𝑖,𝑘 = ℎ 𝑣,𝑖,𝑘 /𝑢 𝑣,𝑖,𝑘 inside the 𝑘-th segment. We de˜ fine the warp, 𝜙 𝑣,𝑖 , for each 𝑥˜ ∈ [𝑋 𝑣,𝑖,𝑘 −1, 𝑋 𝑣,𝑖,𝑘 ] as 𝜙 𝑣,𝑖 (𝑥) = 𝑌𝑣,𝑖,𝑘 −1 + 𝑚 𝑣,𝑖,𝑘 (𝑥˜ − 𝑋 𝑣,𝑖,𝑘 −1 ). Application. Each raw input 𝑥 𝑣 [𝑡, 𝑖] is first min-max normalized to 𝑥˜ 𝑣 [𝑡, 𝑖] ∈ [0, 1] using channel-specific bounds, then the warp is applied to produce the warped value, 𝑥˜ 𝑣′ [𝑡, 𝑖] = 𝜙 𝑣,𝑖 (𝑥˜ 𝑣 [𝑡, 𝑖]). The warp 𝜙 𝑣,𝑖 varies per channel and per patch but is shared across time steps within the input window, so a single learned warp reshapes every channel 𝑣 at patch 𝑖 identically. The warped sequence 𝑥˜ 𝑣′ [𝑡, 𝑖] then forms the input to the temporal modulator (§4.4). Per-patch terrain and foundation-model conditioning. The 2𝐾 · 𝑉 warp parameters per patch are produced by a terrain- and foundationmodel-conditioned module value_modulation, a single convolution layer on the stacked terrain and AEF features followed by a two-layer MLP. We zero-initialize the weights and biases of this module’s final layer, so the softmax over 𝑢 𝑣,𝑖,𝑘 and ℎ 𝑣,𝑖,𝑘 yields uniform widths and heights at the start of training. The initial warp is ˜ = 𝑥, ˜ and training departs from therefore the identity map 𝜙 𝑣,𝑖 (𝑥) this identity as each patch’s terrain and AEF context pushes its breakpoints apart.

4.4

Temporal Modulator

Gamma kernels and two-component mixture. For each channel 𝑣 and lag 𝜏 ∈ {0, . . . , 𝐿−1} with kernel lookback 𝐿 = 48, a single gamma kernel parameterized by peak lag 𝑝 and timescale 𝜃 is the discretized, normalized Gamma(𝛼, 𝜃 ) density: 𝑘 (𝜏; 𝑝, 𝜃 ) = 𝑝 𝜏 𝛼 −1 𝑒 −𝜏 /𝜃 Í𝐿−1 ′ 𝛼 −1 −𝜏 ′ /𝜃 , where 𝛼 = 𝜃 + 1. We reparameterize the shape 𝜏 ′ =0

(𝜏 )

𝑒

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Kawakami et al.

Modality

Source

Shape

Hydrometeorology Terrain Foundation Model Exposure Vulnerability

Precipitation (NOAA AORC [21]); soil moisture and snow water equivalent (NWM v3 [18]) HAND [20], TWI [29], GCN (dry/avg/wet) [33], POLARIS (𝐾𝑠𝑎𝑡 , clay, 𝜃 𝑠 ) [15] AlphaEarth embeddings [13] NSI building inventory [69], NFIP active policies [25] NSI structural [69], NFIP policy attributes [25]

7 × 7 × 73 × 3 56 × 56 × 8 56 × 56 × 64 56 × 56 × 2 56 × 56 × 10

Table 1: DELUGE’s input modalities and tensor shapes (𝐻 × 𝑊 × 𝐶). For hydrometeorology, the third axis is the 𝑇 = 73-step hourly window and the final dimension stacks the three channels; for other modalities the last axis is the feature-channel count.

~8x8km context; 48 hour lookback Precipitation

...

Spatial Context: Terrain and AlphaEarth Conditioning - Same ~8x8km Conv2D Encoder

32 Parameters per variable (16 control points: xi, yi)

Time Snow Water Equivalent

...

...

Value Modulator

Learns a weighting function over each variable’s raw value ranges. Runs variable-wise to normalize each variable to [0,1]

5 Parameters per variable

Temporal Modulator

Learns a temporal lookback weighting function over the 48 hour look back window via a two-mixture Gamma kernel Mixture (normalized) Component 1 Component 2

Model emphasized these value ranges Weight

...

Conv2D Encoder

CapturesS Long-termS Response

CapturesS Short-termS Response

Time

0 0.1 0.2 0.3 0.4 0.5 Variable value, e.g. (Volumetric Soil Moisture (m^3/m^3))

0

12 24 Lookback Time

36 (hours)

Modulated output

Time Soil Moisture

..ˆ

[Height above nearest drainage, Topological Wetness Index,p Global Curve Numbers (Dry, Average, Wet), POLARIS Soil Properties (ksat, clay, theta_s), Google AlphaEarth]

Weight

Input: Hydrometeorological Time Series

48

Figure 3: The two terrain- and AEF-conditioned modulators acting on the hydrometeorology branch. The Value Modulator (§4.3) applies a per-channel, per-patch piecewise-linear warp 𝜙 𝑣,𝑖 : [0, 1] → [0, 1] to each min-max normalized input, reshaping the response curve along the value axis. The Temporal Modulator (§4.4) then convolves the warped sequence with a per-channel gamma-kernel mixture 𝑘 𝑣 (𝜏), integrating the 𝑇 -hour lookback along the time axis. Each modulator has its own conditioning network (ConvBlock + MLP) over the stacked terrain and AlphaEarth Foundation features, making the learned warp shapes and kernel peaks/decay timescales directly interpretable per patch. via 𝛼 = 𝑝/𝜃 + 1 so that the kernel’s peak lag is exactly 𝑝 (the mode of Gamma(𝛼, 𝜃 ) lies at (𝛼 − 1)𝜃 = 𝑝), making the two learned parameters (𝑝, 𝜃 ) directly interpretable: 𝑝 is the peak response time in hours and 𝜃 is the decay timescale in hours. The gamma is chosen here as a flexible, two-parameter kernel often used in hydrology to model rainfall-runoff response [49]. To capture varied temporal patterns, we adopt a per-channel convex mixture of two gammas with the same functional form but different parameters for the precipitation channel. Here, we parameterize a convex mixture of two gamma, weighted by a learnt mixing weight 𝑤 𝑣 : 𝑘 𝑣 (𝜏) = 𝑤 𝑣 𝑘 (𝜏; 𝑝 𝑣1, 𝜃 𝑣1 ) + (1 − 𝑤 𝑣 ) 𝑘 (𝜏; 𝑝 𝑣2, 𝜃 𝑣2 ). For soil moisture and snow water equivalent, we employ a single gamma. In all, each channel’s temporal response kernel is parameterized by five (precipitation) or two (rest) parameters, totaling kernel parameters per patch.

Per-patch terrain conditioning. The temporal_modulation module is structurally analogous to value_modulation, producing the 9 kernel parameters per patch from the stacked terrain and AEF features. The final linear layer is zero-initialized in weight, with bias chosen so each cell decodes the same hydrologically reasonable kernel at step 0 (𝑝 1 ≈ 3h, 𝑝 2 ≈ 24h, 𝜃 1 ≈ 3h, 𝜃 2 ≈ 12h, 𝑤 ≈ 0.5); training then pushes per-patch parameters apart based on each patch’s characteristics.

Causal convolution. Let 𝑖 be one of the pixels within the 7 × 7 patch, and let 𝑥 𝑣 [𝑡, 𝑖] denote the value of hydrometeorology channel 𝑣 at time 𝑡 and pixel 𝑖 within the 𝑇 -hour input window. The kernel 𝑘 𝑣 is then convolved causally with the input sequence to produce Í𝐿−1 the response at pixel 𝑖: 𝑟 𝑣 [𝑡, 𝑖] = 𝜏=0 𝑘 𝑣 (𝜏) 𝑥 𝑣 [𝑡 −𝜏, 𝑖]. The kernel is shared across all pixels within the patch. Temporal reduction. The temporal axis is collapsed by selecting one time step 𝑡 ′ per pixel, the moment of peak accumulated precipitation response within Day 𝐷: 𝑡 ′ (𝑖) = argmax𝑡 ∈Day D 𝑟 precip [𝑡, 𝑖]. We reduce this via, 𝑟 reduced [𝑣, 𝑖] = 𝑟 𝑣 [𝑡 ′ (𝑖), 𝑖]. All channels are read at the same 𝑡 ′ , yielding a temporally-coherent Day-𝐷 snapshot. We expect this to summarize the hydrometeorological state at the moment of peak precipitation impact, appropriately weighing the lookback hours according to the learned kernel. Finally, a linear 1×1 convolution lifts the per-pixel 𝑉 -dimensional reduced response to a 32D feature space for fusion with the other modalities.

4.5

Other Components

Modality Encoders. The remaining three modalities are processed by lightweight convolutional encoders, all operating on the same ∼8 × 8 km patch. The NSI building-inventory and structural fields are native to the 56 × 56 grid, while the NFIP active-policy and policy-attribute fields are native to the 7 × 7 grid and are bilinearly upsampled to 56 × 56, so the Terrain-Hazard, Exposure, and Vulnerability encoders all operate at a common 56 × 56 resolution. Each

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

encoder maps its input through strided convolutional blocks that downsample to 7 × 7, producing 16, 8, and 8 channels respectively. All encoders use 3 × 3 convolutions, leaky-ReLU activations, and batch normalization [32]. Cross-Modal Fusion. At each of the 7×7 spatial positions, the four modality feature vectors form a set of four tokens. Each token is linearly projected to 𝑑 model = 32 by a 1 × 1 convolution and summed with a learnable modality-type embedding that identifies its source modality. We fuse the tokens with a single Transformer encoder layer [71], applying 4-head self-attention and a position-wise feedforward network, each wrapped in a residual connection and layer normalization, so that every modality attends to the others before fusion. The fused representation 𝑓 is read out at the center pixel to align with the per-pixel prediction target. Classifier. The center-pixel feature vector 𝑓 ∈ R𝑑model is mapped to a claim probability by an MLP classification head with a single hidden layer. The hidden layer applies layer normalization and a GELU nonlinearity [27], and a final linear layer with a sigmoid activation produces the per-pixel probability 𝑦ˆ ∈ [0, 1].

4.6

Training

Spatial Holdout Split. To prevent spatial leakage, we partition the 100 grid cells (§3.2) into train and test sets at the 75 km cell level rather than at the pixel level. We use a spatially blocked split of grid cells, so the model is evaluated on held-out geographic regions rather than on held-out samples from the same regions. Because our 100 grid cells span the diverse hydroclimatic regions of CONUS, across which flood dynamics are known to vary significantly [14], evaluating on spatially held-out cells tests whether DELUGE generalizes across this wide range of flood dynamics rather than to new samples from cells it has already seen. Loss. We use focal loss [42] with 𝛼 = 0.5, 𝛾 = 2, weighted per sample by our claim polygon area based confidence score 𝑐 (§3.3). This choice is primarily motivated by the extreme class imbalance and to de-emphasize the easy negative samples, which are abundant in the dataset and can overwhelm the learning signal from the more informative positive samples. In addition to the main loss on the predicted flood probability, we add a smoothness penalty on the Value Modulator’s parameters Í Í𝐾 −1 (§4.3): Lsmooth = 𝑣 𝐾 1−1 𝑘=1 (𝑢 𝑣,𝑘+1 − 𝑢 𝑣,𝑘 ) 2 + (ℎ 𝑣,𝑘+1 − ℎ 𝑣,𝑘 ) 2 . The penalty encourages adjacent widths and adjacent heights to vary smoothly across segments, biasing the learned warp toward a smooth function. In all, we optimize for the weighted sum of the focal loss and the smoothness penalty: L = Lfocal + 𝜆 · Lsmooth with 𝜆 = 0.1. We apply the AdamW optimizer [44] with a cosine learning-rate schedule and linear warmup.

5

Results

To evaluate the performance of DELUGE, we compare it against tree-based baselines, specifically Random Forest, XGBoost [16], and LightGBM [36]. We focus on tree-based baselines for two reasons. First, the existing ML literature on flood-damage prediction overwhelmingly relies on tabular models, with Random Forest the most commonly reported baseline [2, 17, 26, 41, 76]; we therefore include Random Forest alongside XGBoost and LightGBM as stronger

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

tuned variants of the same family. Second, while graph neural network (GNN) approaches have been proposed for flood prediction [35, 61, 66], they have been developed primarily for fluvial dynamics, predicting inundation depth, and demonstrated on small catchment-scale domains, making them impractical for a CONUSwide pluvial setting. Evaluation Metrics. Given the severe class imbalance of our problem, we use the Precision-Recall AUC (PR-AUC) as our evaluation metric. In addition, we develop a dollar-weighted version of PRAUC to account for varied range in reported flood damages. To this end, we associate each positive sample with the damage density where the claim damage is equally distributed across the pixels in the claim’s polygon footprint. Then, integrating under a precisionrecall curve where the recall is defined as the percentage of total claim dollars correctly predicted as flood claims, with precision as its standard definition, we obtain a dollar-weighted PR-AUC metric that captures the model’s ability to identify the claims weighted by claimed damages. Under this metric, we evaluate the performance of our model in terms of its ability to capture the most costly flood claims, which is of interest to insurance companies and other stakeholders.

5.1

Comparison against Baselines

Table 2 presents the headline performance comparison between DELUGE and baseline models. Since tree-based models do not natively handle time-series or spatial rasters, we use the same set of features as introduced in §4.2, but with some modifications to make them compatible with tree-based models. For the hydrometeorological time-series, we extract the daily sum and daily max for each of the hydrometeorological features as often done in literature [2, 76] and the center pixel value and the spatial mean for all other features in the given patch. We tuned the tree-based baselines (RF, XGBoost, LightGBM) by randomized search over 50 configurations. DELUGE outperforms all three baselines on both metrics, with LightGBM the strongest tree baseline (Table 2). Because every method is evaluated on the same splits, Δ% is a mean per-seed difference rather than a ratio of marginal means. DELUGE leads every baseline on this paired basis, so the overlapping absoluteerror bars reflect shared seed-to-seed variance, not ambiguity in the ranking. The relative gap widens on the dollar-weighted metric, with the tree baselines trailing DELUGE by ∼9 to 30% on $-Weighted PR-AUC versus ∼6 to 27% on PR-AUC. This pattern suggests that DELUGE’s architecture helps capture the high-cost claims that tabular baselines under-predict, the regime most relevant to insurance and emergency-management stakeholders. Figure 4 shows observed claims alongside DELUGE predictions drawn from the held-out spatial split. Operating-Point Analysis. To probe how this dollar concentration manifests at fixed alert budgets, we report top-𝐾 recall at several values of 𝐾 (Table 3), separating occurrence recall (R@K) from dollar recall ($R@K). At every 𝐾, $R@K sharply exceeds R@K. At 𝐾 = 0.1% DELUGE captures only 19% of pluvial claim events but 62% of total damage dollars and at 𝐾 = 1% the gap remains (52% events vs. 90% dollars). DELUGE also leads both XGBoost and LightGBM at every 𝐾 on both metrics, with the largest gap at the tightest, most operationally relevant budgets. DELUGE’s

Kawakami et al.

Predicted

NFIP Claims

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Figure 4: Pluvial NFIP claims (top) and DELUGE predictions (bottom) on seven held-out test days. Red colored cells corresponds to pixels with flood claims (top) and pixel where DELUGE predicted pluvial NFIP flood claims (bottom). Model

PR-AUC

Δ%

$-W PR-AUC

Δ%

DELUGE (Ours) Random Forest XGBoost [16] LightGBM [36]

0.243±0.007 0.177±0.014 0.219±0.006 0.228±0.015

– −27% −10% −6%

0.594±0.049 0.413±0.016 0.511±0.018 0.538±0.030

– −30% −14% −9%

Table 2: Comparison of DELUGE against tree-based baselines. PR-AUC measures binary damage-occurrence prediction, and $-Weighted PR-AUC weights predictions by claim amount. At a ∼0.25% prevalence, DELUGE shows a ∼100x improvement over chance performance. Scores are reported over three matched spatial splits, and Δ% is each baseline’s mean per-seed relative change from DELUGE. top-ranked predictions are therefore concentrated in the high-cost regime, the operationally meaningful slice of the test set for insurance and emergency-management workflows. DELUGE

XGBoost

LightGBM

𝑲

R@K

$R@K

R@K

$R@K

R@K

$R@K

0.1% 0.2% 0.5% 1%

0.193 0.279 0.412 0.521

0.617 0.728 0.832 0.897

0.173 0.257 0.388 0.495

0.609 0.701 0.813 0.880

0.180 0.260 0.391 0.500

0.557 0.668 0.803 0.872

Table 3: Top-𝐾 recall. R@K is occurrence recall at 𝐾 (fraction of positive pixel-days captured) and $R@K is dollar recall at 𝐾 (fraction of total claim dollars captured). 𝐾 given as a percentile of the held-out pixel-day test set.

5.2

Ablation Studies

We ablate DELUGE along two axes (Table 4). Module ablations swap out our architectural contributions and feature ablations drop one input modality at a time while leaving the architecture intact. Modulator Ablations. We replace our modulator-based encoder with a standard ConvLSTM and a Transformer, common architectures for spatiotemporal data, applied to the hydrometeorological time-series. DELUGE performs on par with both alternatives on mean, with the Transformer marginally ahead on PR-AUC and DELUGE marginally ahead on $-Weighted PR-AUC, while ConvLSTM trails slightly on both. Our modulator design therefore retains predictive accuracy comparable to stronger sequence-model

alternatives while adding interpretability-by-design. We then progressively ablate the modulators’ conditioning, from frozen at initialization (Uniform), to learnable but globally shared, to perpatch conditioned on a single channel (AEF-only or Terrain-only). Each step closes part of the gap to full DELUGE on both metrics, confirming that both learning the modulator parameters and conditioning them per patch matter. AEF-only and Terrain-only perform nearly identically on PR-AUC, with AEF-only modestly stronger on $-Weighted PR-AUC, suggesting the two conditioning channels carry largely overlapping information about local hydrology. This demonstrates the utility of AlphaEarth embeddings for pluvial flood prediction. As a conditioning signal they perform on par with a curated suite of hydrological terrain descriptors and well above unconditioned modulators, with no hydrological feature engineering required. Feature Ablations. The single largest drop comes from removing the hydrometeorological hazard inputs (precipitation, soil moisture, and snow water equivalent), which collapses the model and confirms hydrometeorology as the dominant driver of DELUGE’s prediction. Among the remaining modalities, Exposure shows the largest gap, reflecting its role in identifying which structures sit in the path of damaging precipitation. Vulnerability shows the smallest gap, suggesting structural attributes carry less independent signal, perhaps as they correlate with AlphaEarth [8].

5.3

Verifying Learned Behavior of the Modulators

Explainability and physical fidelity are growing concerns in Geospatial AI [30, 59, 64] and flood modeling [66]. While a substantial body of work examines what geospatial embeddings encode about physical space [9, 55], comparatively little addresses how downstream models should consume these representations in a physically verifiable manner. DELUGE’s modulators are designed to this end. Because their parameters constitute the model’s predictive computation directly, the interpretability they afford is architectural rather than post-hoc, obtained by inspecting the learned parameters themselves rather than by explaining individual predictions. The modulators act on the hydrometeorology branch, which our ablations also identify as the dominant predictive driver (Table 4), so its learned response is the most consequential to interpret. For ˜ and each patch, we summarize the Value Modulator’s warp 𝜙 𝑣,𝑖 (𝑥)

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Figure 5: Learned modulator dynamics clustering result for five representative held-out cells (§5.3). Colors follows the same scheme as Fig. 6 for consistency.

0.243±0.007

Module ablations Uniform Modulators Globally-shared Modulators AEF-only Modulators Terrain-only Modulators

0.196±0.014 −19% 0.217±0.011 −11% 0.234±0.016 −4% 0.235±0.012 −3%

0.594±0.049

0.414±0.050 −30% 0.528±0.053 −11% 0.572±0.081 −4% 0.558±0.073 −6%

ConvLSTM Hydromet. Transformer Hydromet.

0.234±0.014 0.250±0.029

−4% +3%

0.543±0.089 0.574±0.119

−9% −4%

Feature ablations Exposure Vulnerability Terrain Hazard Hydromet. Hazard

0.217±0.017 −11% 0.229±0.012 −6% 0.218±0.007 −11% 0.020±0.002 −92%

0.476±0.071 0.526±0.056 0.496±0.078 0.038±0.008

−20% −11% −17% −94%

Table 4: Module and feature ablations of DELUGE. Each row removes or replaces the listed component from the full model. Module ablations swap an architectural component while feature ablations drop one input modality while leaving the architecture intact. Temporal Modulator’s response kernel 𝑘 𝑣 (𝜏) by their CDFs. Here, we omit snow water equivalent and focus on precipitation and soil moisture, the channels most directly relevant to pluvial flooding. To cluster, we apply PCA on the CDFs down to two components (∼75% of variance) followed by K-means. A silhouette sweep determined 𝐾 = 4. To illustrate the learned modulator behavior within each cluster, Figure 5 shows representative held-out grid cells with their modulator dynamics clustering assignments, and Figure 6 shows the corresponding modulator dynamics for each cluster. First, examining the Temporal Modulator’s kernel (Fig. 6) for 1hr acc. precip. (top) and soil moisture (bottom), we find that DELUGE learns to rely on precipitation for the immediate risk and soil moisture to capture antecedent conditions, with the latter kernels peaking in the ∼7 − 15 hour range in contrast with the precipitation kernels that peak at 𝜏 ≈ 0 and decay rapidly. Notice in Fig 6 that within every cluster the two kernels are cleanly staggered, with the precipitation kernel concentrated at the most recent hour and the soil-moisture kernel peaking several hours later, an explicit hand-off from instantaneous precipitation forcing to antecedent wetness state that DELUGE recovers without a hydrological prior. Next, with a further qualitative inspection of the clustering results, we derive the characters of the four identified clusters. Cluster 1 (Blue) corresponds to denser urban areas, with a sharp reliance on recent precipitation. This short lag time aligns with

Value Modulator

Temporal Modulator Kernel Weight

DELUGE (Full)

Δ%

0.10 0.05

0

20

40

60

80

1hr Acc. Precip. Value

100

Kernel Weight

Δ% $-W PR-AUC

1hr Acc. Precip. Warp Weight

PR-AUC

Soil Moisture Warp Weight

Configuration

0.10 0.05 0.00

0.0

0.1

0.2

0.3

0.4

Soil Moisture Value

0.5

Cluster 1 Cluster 2 Cluster 3 Cluster 4

0.2 0.1 0.0

0

10

20

30

40

0

10

20

30

40

0.15 0.10 0.05 0.00

Lookback τ (hours)

Figure 6: Clusters identified in modulator-parameter space (𝐾 = 4) and their corresponding modulator dynamics. how impervious surfaces accelerate runoff yield [41]. Furthermore, it shows higher importance of lower precipitation than others, signaling its relative vulnerability to lower precipitation and relies on relatively recent soil moisture with emphasis on the higher soil moisture values. Cluster 2 (Orange) corresponds to suburban areas in inland regions of CONUS, with moderate reliance on recent precipitation. Soil moisture kernel is similar to cluster 1, with a slightly even emphasis on all soil moisture values. Cluster 3 (Green) corresponds to suburban areas along the Gulf, with a longer tail of precipitation kernel and a longer and wider lookback range than the first two clusters. Its characteristics lie in the middle of the clusters 2 and 4 drawing on both the suburban and water-facing characteristics. Cluster 4 (Red) corresponds to water-facing regions. Represented by the longest tail in the precipitation kernel and the widest lookback range in the soil moisture kernel, this cluster also places the most even importance in soil moisture values. This may suggest water-facing regions are more susceptible to flooding even when the top layer of soil is not saturated due to the high water table. We emphasize that the value of this result lies not in the clusters aligning with geography. Such alignment is unsurprising, since the modulators are conditioned on foundation-model embeddings and terrain descriptors that already encode geographic structure. It lies instead in the fact that the clusters admit a coherent hydrological interpretation, the architecture-level interpretability we set out to establish, which allows the physical fidelity of the foundation-model conditioning to be assessed directly from the learned parameters. More broadly, we view this conditioning scheme as not unique to pluvial flooding but applicable to

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Kawakami et al.

Predicted

NFIP Claims

6

Figure 7: DELUGE dense urban failure cases NFIP claims and DELUGE predictions on four held-out test days where DELUGE fails to pinpoint flood claims in urban areas. other geospatial tasks that build on Geospatial Foundation Models, where routing embeddings through physically meaningful parametric modules can make the model’s use of them directly inspectable. The full clustering results are placed in the Appendix.

5.4

Failure Modes

Low-damage claims are harder to predict. These claims are often associated with less anomalous hydrometeorological conditions, making them more difficult to distinguish from non-damage cases. To characterize this quantitatively, we stratify held-out claim-days by claim amount (Table 5). PR-AUC rises sharply from the smallest bin to the [$50K, $500K) tier (0.04 → 0.37), with $-Weighted PR-AUC tracking the same trend. The slight drop in the largest [$500K, ∞) tier reflects its extreme base rate (0.004%) rather than degraded ranking. DELUGE is therefore strongest in the highcost regime most relevant to insurance and emergency management stakeholders. Damage range ($) [0, 5,000) [5,000, 50,000) [50,000, 500,000) [500,000, ∞)

Prev. (%)

PR-AUC

$-W PR-AUC

0.124 0.087 0.030 0.004

0.043 0.154 0.368 0.313

0.049 0.186 0.424 0.331

Table 5: DELUGE performance stratified by claim amount. Per-bin prevalence, PR-AUC, and dollar-weighted PR-AUC on a single held-out spatial split. Locations of impacts of large storm systems with widespread precipitation are hard to resolve. Pinpointing which pixels experience damage is difficult, especially in dense urban areas where flood response is governed by fine-scale local processes. Figure 7 shows representative cases where DELUGE registers that a large storm system causes damage but struggles to localize exactly which urban pixels are affected. We attribute this to two factors. First, because we predict only insured damage reported through NFIP, some pixels in our negative class may have experienced unreported damage, blurring the supervisory signal. Second, and independent of this, dense urban areas are where we most consistently observe DELUGE to struggle, as we show by stratifying performance over AEF-defined location characteristics in the Appendix. We anticipate that higher spatial (sub-kilometer) and temporal (sub-hourly) resolution hydrometeorology will be needed to resolve fine-scale urban damage.

Discussion and Conclusion

We presented DELUGE, a multimodal CNN-based system for daily pluvial insured flood damage prediction at ∼1 km resolution across the highest-claim regions of CONUS, trained on NFIP claims. We structured the model around the core disaster-management components of hazard, exposure, and vulnerability, introduced an interpretable conditioning scheme for integrating foundation-model embeddings into the prediction task, and inspected the physical fidelity of that scheme. As a prerequisite for using NFIP claims as a daily training signal, we also introduced precipitation-guided temporal and intersection-based spatial corrections that refine the timing and footprint of each claim. DELUGE outperforms tuned gradient-boosted tree baselines on PR-AUC and concentrates highdollar claims more effectively near the top of its prediction ranking, the regime most relevant to insurance and emergency-management stakeholders. For insurers, these gains can improve risk segmentation. When calibrated to average annual loss, a model that better identifies high-dollar pluvial flood risk could support territory rating, risk selection, and reinsurance analysis, while also providing clearer risk signals for policyholders and communities. We further demonstrated the efficacy of our Value and Temporal Modulators, which gain interpretability-by-design while remaining competitive with standard encoders for spatial time series. We believe this offers a scheme transferable to other geospatial tasks built on nascent Geospatial Foundation Models. Three limitations point to natural next steps. First, data quality remains the dominant bottleneck, on both the label and input sides. On the label side, NFIP claims, despite our spatial and temporal uncertainty corrections, remain a noisy and incomplete proxy for pluvial damage. Uninsured losses and unreported events are systematically absent from the record, and claimed damage amount is itself an imperfect proxy for severity or societal importance. The resulting sparsity in both space and time is reflected in the modest absolute performance of all models including DELUGE, which nonetheless remains strong when the low prevalence of claims is taken into account. On the input side, our hydrometeorology inputs themselves carry nontrivial uncertainty at the hourly kilometer scale. Improving the fidelity of both the damage signal and the hydrometeorological hazard inputs is likely the single largest lever for further predictive gains. Second, interpretability remains open. Our conditioning scheme is a step forward in that it exposes how the model utilizes foundation-model embeddings, but it does not explain individual predictions. Closing this gap, from architecturelevel fidelity to case-level explanation, is a natural direction for future work. Third, our spatial evaluation shows that DELUGE generalizes across the wide range of flood dynamics present within highest-claim regions in CONUS [14], but it does not yet test extrapolation to geographies absent from training. Applying DELUGE to regions or countries outside its training distribution, where flood dynamics and the built environment may differ from anything it has seen, is a distinct problem that we leave to future work. Together, these results suggest that DELUGE can serve as a starting point for CONUS-scale pluvial flood impact modeling, and that interpretable deep learning can turn geospatial foundation-model embeddings into accurate and physically meaningful prediction systems.

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

References [1] Mohit Agarwal, Mimi Sun, Chaitanya Kamath, Arbaaz Muslim, Prithul Sarker, Joydeep Paul, Hector Yee, Marcin Sieniek, Kim Jablonski, Swapnil Vispute, Atul Kumar, Yael Mayer, David Fork, Sheila de Guia, Jamie McPike, Adam Boulanger, Tomer Shekel, David Schottlander, Yao Xiao, Manjit Chakravarthy Manukonda, Yun Liu, Neslihan Bulut, Sami Abu el haija, Bryan Perozzi, Monica Bharel, Von Nguyen, Luke Barrington, Niv Efron, Yossi Matias, Greg Corrado, Krish Eswaran, Shruthi Prabhakara, Shravya Shetty, and Gautam Prasad. 2026. General Geospatial Inference with a Population Dynamics Foundation Model. (2026). arXiv:2411.07207 [cs.LG] https://arxiv.org/abs/2411.07207 [2] Atieh Alipour, Ali Ahmadalipour, Peyman Abbaszadeh, and Hamid Moradkhani. 2020. Leveraging machine learning for predicting flash flood damage in the Southeast US. Environmental Research Letters 15, 2 (Feb. 2020), 024011. doi:10. 1088/1748-9326/ab6edd [3] American Meteorological Society. 2026. Glossary of Meteorology: Rain. https: //glossary.ametsoc.org/wiki/rain Accessed: 2026-06-09. [4] Natee Amornsiripanitch, Siddhartha Biswas, John Orellana, and David Zink. 2024. Flood Underinsurance. (2024). [5] P.D Bates and A.P.J De Roo. 2000. A simple raster-based model for flood inundation simulation. Journal of Hydrology 236, 1-2 (Sept. 2000), 54–77. doi:10.1016/S0022-1694(00)00278-X [6] Paul D. Bates. 2022. Flood Inundation Prediction. Annual Review of Fluid Mechanics 54, Volume 54, 2022 (Jan. 2022), 287–315. doi:10.1146/annurev-fluid-030121113138 [7] Paul D. Bates, Niall Quinn, Christopher Sampson, Andrew Smith, Oliver Wing, Jeison Sosa, James Savage, Gaia Olcese, Jeff Neal, Guy Schumann, Laura Giustarini, Gemma Coxon, Jeremy R. Porter, Mike F. Amodeo, Ziyan Chu, Sharai Lewis-Gruss, Neil B. Freeman, Trevor Houser, Michael Delgado, Ali Hamidi, Ian Bolliger, Kelly E. McCusker, Kerry Emanuel, Celso M. Ferreira, Arslaan Khalid, Ivan D. Haigh, Anaïs Couasnon, Robert E. Kopp, Solomon Hsiang, and Witold F. Krajewski. 2021. Combined Modeling of US Fluvial, Pluvial, and Coastal Flood Hazard Under Current and Future Climates. Water Resources Research 57, 2 (2021), e2020WR028673. doi:10.1029/2020WR028673 _eprint: https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1029/2020WR028673. [8] Aaron Bell, Amit Aides, Amr Helmy, Arbaaz Muslim, Aviad Barzilai, Aviv Slobodkin, Bolous Jaber, David Schottlander, George Leifman, Joydeep Paul, Mimi Sun, Nadav Sherman, Natalie Williams, Per Bjornsson, Roy Lee, Ruth Alcantara, Thomas Turnbull, Tomer Shekel, Vered Silverman, Yotam Gigi, Adam Boulanger, Alex Ottenwess, Ali Ahmadalipour, Anna Carter, Behzad Vahedi, Charles Elliott, David Andre, Elad Aharoni, Gia Jung, Hassler Thurston, Jacob Bien, Jamie McPike, Jessica Sapick, Juliet Rothenberg, Kartik Hegde, Kel Markert, Kim Philipp Jablonski, Luc Houriez, Monica Bharel, Phing VanLee, Reuven Sayag, Sebastian Pilarski, Shelley Cazares, Shlomi Pasternak, Siduo Jiang, Thomas Colthurst, Yang Chen, Yehonathan Refael, Yochai Blau, Yuval Carny, Yael Maguire, Avinatan Hassidim, James Manyika, Tim Thelin, Genady Beryozkin, Gautam Prasad, Luke Barrington, Yossi Matias, Niv Efron, and Shravya Shetty. 2026. Earth AI: Unlocking Geospatial Insights with Foundation Models and Cross-Modal Reasoning. doi:10.48550/arXiv.2510.18318 arXiv:2510.18318 [cs.AI]. [9] Ivan Felipe Benavides-Martinez, Justin Guthrie, Jhon Edwin Arias, Yeison Alberto Garces-Gomez, Angela Ines Guzman-Alvis, Cristiam Victoriano Portilla-Cabrera, Somnath Mondal, Andrew J Allyn, and Auroop R Ganguly. 2026. What on Earth is AlphaEarth? Hierarchical structure and functional interpretability for global land cover. arXiv preprint arXiv:2603.16911 (2026). [10] Derrick Bonafilia, Beth Tellman, Tyler Anderson, and Erica Issenberg. 2020. Sen1Floods11: a georeferenced dataset to train and test deep learning flood algorithms for Sentinel-1. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 835–845. doi:10.1109/CVPRW50498. 2020.00113 [11] Anasse Boutayeb, Iyad Lahsen-Cherif, and Ahmed El Khadimi. 2025. When Machine Learning Meets Geospatial Data: A Comprehensive GeoAI Review. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 18 (2025), 13135–13191. doi:10.1109/JSTARS.2025.3568715 [12] C. Bouwens, M.-C. ten Veldhuis, M. Schleiss, X. Tian, and J. Schepers. 2018. Towards identification of critical rainfall thresholds for urban pluvial flooding prediction based on crowdsourced flood observations. Hydrology and Earth System Sciences Discussions 2018 (2018), 1–24. doi:10.5194/hess-2017-751 [13] Christopher F. Brown, Michal R. Kazmierski, Valerie J. Pasquarella, William J. Rucklidge, Masha Samsikova, Chenhui Zhang, Evan Shelhamer, Estefania Lahera, Olivia Wiles, Simon Ilyushchenko, Noel Gorelick, Lihui Lydia Zhang, Sophia Alj, Emily Schechter, Sean Askay, Oliver Guinan, Rebecca Moore, Alexis Boukouvalas, and Pushmeet Kohli. 2025. AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data. doi:10.48550/ arXiv.2507.22291 arXiv:2507.22291 [cs.CV]. [14] Manuela I. Brunner, Eric Gilleland, Andy Wood, Daniel L. Swain, and Martyn Clark. 2020. Spatial Dependence of Floods Shaped by Spatiotemporal Variations in Meteorological and Land-Surface Processes. Geophysical Research Letters 47, 13 (2020), e2020GL088000. doi:10.1029/2020GL088000

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

[15] Nathaniel W. Chaney, Budiman Minasny, Jonathan D. Herman, Travis W. Nauman, Colby W. Brungard, Cristine L. S. Morgan, Alexander B. McBratney, Eric F. Wood, and Yohannes Yimam. 2019. POLARIS Soil Properties: 30-m Probabilistic Maps of Soil Properties Over the Contiguous United States. Water Resources Research 55, 4 (2019), 2916–2938. arXiv:https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1029/2018WR022797 doi:10.1029/2018WR022797 [16] Tianqi Chen and Carlos Guestrin. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (San Francisco, California, USA) (KDD ’16). Association for Computing Machinery, New York, NY, USA, 785–794. doi:10. 1145/2939672.2939785 [17] Elyssa L Collins, Georgina M Sanchez, Adam Terando, Charles C Stillwell, Helena Mitasova, Antonia Sebastian, and Ross K Meentemeyer. 2022. Predicting flood damage probability across the conterminous United States. Environmental Research Letters 17, 3 (Feb. 2022), 034006. doi:10.1088/1748-9326/ac4f0f [18] Brian Cosgrove, David Gochis, Trey Flowers, Aubrey Dugger, Fred Ogden, Tom Graziano, Ed Clark, Ryan Cabell, Nick Casiday, Zhengtao Cui, Kelley Eicher, Greg Fall, Xia Feng, Katelyn Fitzgerald, Nels Frazier, Camaron George, Rich Gibbs, Liliana Hernandez, Donald Johnson, Ryan Jones, Logan Karsten, Henok Kefelegn, David Kitzmiller, Haksu Lee, Yuqiong Liu, Hassan Mashriqui, David Mattern, Alyssa McCluskey, James L. McCreight, Rachel McDaniel, Alemayehu Midekisa, Andy Newman, Linlin Pan, Cham Pham, Arezoo RafieeiNasab, Roy Rasmussen, Laura Read, Mehdi Rezaeianzadeh, Fernando Salas, Dina Sang, Kevin Sampson, Tim Schneider, Qi Shi, Gautam Sood, Andy Wood, Wanru Wu, David Yates, Wei Yu, and Yongxin Zhang. 2024. NOAA’s National Water Model: Advancing operational hydrology through continental-scale modeling. JAWRA Journal of the American Water Resources Association 60, 2 (2024), 247–272. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/1752-1688.13184 doi:10.1111/1752-1688.13184 [19] Elena Cristiano, Marie-Claire ten Veldhuis, and Nick van de Giesen. 2017. Spatial and temporal variability of rainfall and their effects on hydrological response in urban areas – a review. Hydrology and Earth System Sciences 21, 7 (July 2017), 3859–3878. doi:10.5194/hess-21-3859-2017 [20] European Space Agency and Sinergise. 2021. Copernicus Digital Elevation Model (DEM). Registry of Open Data on AWS. https://registry.opendata.aws/copernicusdem/ Accessed: 2026-05-18. [21] Greg Fall, David Kitzmiller, Sandra Pavlovic, Ziya Zhang, Nathan Patrick, Michael St. Laurent, Carl Trypaluk, Wanru Wu, and Dennis Miller. 2023. The Office of Water Prediction’s Analysis of Record for Calibration, version 1.1: Dataset description and precipitation evaluation. JAWRA Journal of the American Water Resources Association 59, 6 (2023), 1246–1272. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/1752-1688.13143 doi:10. 1111/1752-1688.13143 [22] Federal Emergency Management Agency. 2021. FEMA Flood Zone. FEMA-NFHL. https://msc.fema.gov Publication Date: December 26, 2021. [23] Federal Emergency Management Agency. 2024. The Cost of Flooding. https: //www.floodsmart.gov/know-your-risk/cost-of-flooding. Accessed: 2026-06-19. [24] Federal Emergency Management Agency. 2026. OpenFEMA Dataset: FIMA NFIP Redacted Claims - v2. https://www.fema.gov/openfema-data-page/fima-nfipredacted-claims-v2. [25] Federal Emergency Management Agency. 2026. OpenFEMA Dataset: FIMA NFIP Redacted Policies - v2. https://www.fema.gov/openfema-data-page/fima-nfipredacted-policies-v2. [26] Helena M. Garcia, Antonia Sebastian, Kieran P. Fitzmaurice, Miyuki Hino, Elyssa L. Collins, and Gregory W. Characklis. 2025. Reconstructing Repetitive Flood Exposure Across 78 Events From 1996 to 2020 in North Carolina, USA. Earth’s Future 13, 7 (July 2025), e2025EF006026. doi:10.1029/2025EF006026 [27] Dan Hendrycks and Kevin Gimpel. 2016. Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units. CoRR abs/1606.08415 (2016). arXiv:1606.08415 http://arxiv.org/abs/1606.08415 [28] Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, Adrian Simmons, Cornel Soci, Saleh Abdalla, Xavier Abellan, Gianpaolo Balsamo, Peter Bechtold, Gionata Biavati, Jean Bidlot, Massimo Bonavita, Giovanna De Chiara, Per Dahlgren, Dick Dee, Michail Diamantakis, Rossana Dragani, Johannes Flemming, Richard Forbes, Manuel Fuentes, Alan Geer, Leo Haimberger, Sean Healy, Robin J. Hogan, Elías Hólm, Marta Janisková, Sarah Keeley, Patrick Laloyaux, Philippe Lopez, Cristina Lupu, Gabor Radnoti, Patricia de Rosnay, Iryna Rozum, Freja Vamborg, Sebastien Villaume, and Jean-Noël Thépaut. 2020. The ERA5 global reanalysis. Quarterly Journal of the Royal Meteorological Society 146, 730 (2020), 1999–2049. arXiv:https://rmets.onlinelibrary.wiley.com/doi/pdf/10.1002/qj.3803 doi:10.1002/ qj.3803 [29] Zachary H. Hoylman. 2021. A 30m Topographic Wetness Index Dataset for the Continental United States. doi:10.5281/zenodo.4460354 [30] Chia-Yu Hsu and Wenwen Li. 2023. Explainable GeoAI: can saliency maps help interpret artificial intelligence’s learning process? An empirical study on natural

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

feature detection. International Journal of Geographical Information Science 37, 5 (May 2023), 963–987. doi:10.1080/13658816.2023.2191256 [31] George Huffman, David Bolvin, Dan Braithwaite, Kuolin Hsu, Robert Joyce, Chris Kidd, Eric Nelkin, Soroosh Sorooshian, Erich Stocker, Jackson Tan, and David Wolff. 2020. Integrated Multi-satellite Retrievals for the Global Precipitation Measurement (GPM) Mission (IMERG). 343–353. doi:10.1007/978-3-030-245689_19 [32] Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd International Conference on Machine Learning - Volume 37 (Lille, France) (ICML’15). JMLR.org, 448–456. [33] Hadi H. Jaafar, Farah A. Ahmad, and Naji El Beyrouthy. 2019. GCN250, new global gridded curve numbers for hydrologic modeling and design. Scientific Data 6, 1 (Aug. 2019), 145. doi:10.1038/s41597-019-0155-x [34] Eric P. James and Russ S. Schumacher. 2024. Precipitation Proxies for Flash Flooding: A Seven-Year Analysis over the Contiguous United States. Journal of Hydrometeorology 25, 9 (Sept. 2024), 1323–1344. doi:10.1175/JHM-D-23-0203.1 [35] Arnold Kazadi, James Doss-Gollin, Antonia Sebastian, and Arlei Silva. 2024. FloodGNN-GRU: a spatio-temporal graph neural network for flood prediction. Environmental Data Science 3 (2024), e21. doi:10.1017/eds.2024.19 [36] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. LightGBM: a highly efficient gradient boosting decision tree. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 3149–3157. [37] Vijendra Kumar, Kul Vaibhav Sharma, Tommaso Caloiero, Darshan J. Mehta, and Karan Singh. 2023. Comprehensive Overview of Flood Modeling Approaches: A Review of Recent Advances. Hydrology 10, 7 (July 2023), 141. doi:10.3390/ hydrology10070141 [38] Hyunho Lee and Wenwen Li. 2024. Improving interpretability of deep active learning for flood inundation mapping through class ambiguity indices using multi-spectral satellite imagery. Remote Sensing of Environment 309 (Aug. 2024). doi:10.1016/j.rse.2024.114213 [39] Wenzhong Li, Chengshuai Liu, Yingying Xu, Chaojie Niu, Runxi Li, Ming Li, Caihong Hu, and Lu Tian. 2024. An interpretable hybrid deep learning model for flood forecasting based on Transformer and LSTM. Journal of Hydrology: Regional Studies 54 (Aug. 2024), 101873. doi:10.1016/j.ejrh.2024.101873 [40] Zhi Li, Shang Gao, Mengye Chen, Jonathan J. Gourley, Changhai Liu, Andreas F. Prein, and Yang Hong. 2022. The conterminous United States are projected to become more prone to flash floods in a high-end emissions scenario. Communications Earth & Environment 3, 1 (April 2022), 86. doi:10.1038/s43247-022-00409-6 [41] Yaoxing Liao, Zhaoli Wang, Xiaohong Chen, and Chengguang Lai. 2023. Fast simulation and prediction of urban pluvial floods using a deep convolutional neural network model. Journal of Hydrology 624 (2023), 129945. doi:10.1016/j. jhydrol.2023.129945 [42] Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2020. Focal Loss for Dense Object Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 2 (Feb. 2020), 318–327. doi:10.1109/TPAMI.2018.2858826 [43] Chia-Fu Liu, Lipai Huang, Kai Yin, Sam Brody, and Ali Mostafavi. 2024. FloodDamageCast: Building flood damage nowcasting with machine-learning and data augmentation. International Journal of Disaster Risk Reduction 114 (Nov. 2024), 104971. doi:10.1016/j.ijdrr.2024.104971 [44] Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. doi:10.48550/arXiv.1711.05101 arXiv:1711.05101 [cs.LG]. [45] Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 4768–4777. [46] Roland Löwe and Karsten Arnbjerg-Nielsen. 2020. Urban pluvial flood risk assessment – data resolution and spatial scale when developing screening approaches on the microscale. Natural Hazards and Earth System Sciences 20, 4 (April 2020), 981–997. doi:10.5194/nhess-20-981-2020 [47] Rotem Mayo, Oleg Zlydenko, Moral Bootbool, Shmuel Fronman, Oren Gilon, Avinatan Hassidim, Frederik Kratzert, Gila Loike, Yossi Matias, Yonatan Nakar, Grey Nearing, Reuven Sayag, Amitay Sicherman, Ido Zemach, and Deborah Cohen. 2026. Groundsource: A Dataset of Flood Events from News. doi:10.5281/ ZENODO.18647054 [48] Mohammed Moishin, Ravinesh C. Deo, Ramendra Prasad, Nawin Raj, and Shahab Abdulla. 2021. Designing Deep-Based Learning Flood Forecast Model With ConvLSTM Hybrid Algorithm. IEEE Access 9 (2021), 50982–50993. doi:10.1109/ ACCESS.2021.3065939 [49] J. E. Nash. 1959. Systematic determination of unit hydrograph parameters. Journal of Geophysical Research (1896-1977) 64, 1 (1959), 111–115. arXiv:https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1029/JZ064i001p00111 doi:10.1029/JZ064i001p00111 [50] Grey Nearing, Deborah Cohen, Vusumuzi Dube, Martin Gauch, Oren Gilon, Shaun Harrigan, Avinatan Hassidim, Daniel Klotz, Frederik Kratzert, Asher Metzger, Sella Nevo, Florian Pappenberger, Christel Prudhomme, Guy Shalev, Shlomo

Kawakami et al.

Shenzis, Tadele Yednkachw Tekalign, Dana Weitzner, and Yossi Matias. 2024. Global prediction of extreme floods in ungauged watersheds. Nature 627, 8004 (March 2024), 559–563. doi:10.1038/s41586-024-07145-1 [51] Benjamin Nelson-Mercer, Taeho Kim, Vinh Ngoc Tran, and Valeriy Ivanov. 2025. Pluvial flood impacts and policyholder responses throughout the United States. npj Natural Hazards 2, 1 (Jan. 2025), 8. doi:10.1038/s44304-025-00058-7 [52] Perry C. Oddo, John D. Bolten, Sujay V. Kumar, and Brian Cleary. 2024. Deep Convolutional LSTM for improved flash flood prediction. Frontiers in Water 6 (Feb. 2024). doi:10.3389/frwa.2024.1346104 [53] Ivan Petkov and Francesc Ortega. 2025. Learning from experience: Flooding and insurance take-up in the flood zone and its periphery. Journal of Risk and Insurance 92, 2 (2025), 312–356. doi:10.1111/jori.70002 [54] Supattra Puttinaovarat and Paramate Horkaew. 2020. Flood Forecasting System Based on Integrated Big and Crowdsource Data by Using Machine Learning Techniques. IEEE Access 8 (2020), 5885–5905. doi:10.1109/ACCESS.2019.2963819 [55] Mashrekur Rahman. 2026. Physically interpretable AlphaEarth foundation model embeddings enable LLM-based land surface intelligence. Remote Sensing Applications: Society and Environment 42 (April 2026), 102045. doi:10.1016/j.rsase.2026. 102045 [56] Clément Rambour, Nicolas Audebert, Elise Koeniguer, Bertrand Le Saux, Michel Crucianu, and Mihai Datcu. 2020. SEN12-FLOOD : a SAR and Multispectral Dataset for Flood Detection. doi:10.21227/w6xz-s898 [57] Youtong Rong, Paul Bates, Jeffrey Neal, Leanne Archer, Simbi Hatchard, and Elizabeth Kendon. 2024. Impact of Soil Moisture Dynamics and Precipitation Pattern on UK Urban Pluvial Flood Hazards Under Climate Change. Earth’s Future 12, 5 (2024), e2023EF004073. doi:10.1029/2023EF004073 [58] Bernice R. Rosenzweig, Lauren McPhillips, Heejun Chang, Chingwen Cheng, Claire Welty, Marissa Matsler, David Iwaniec, and Cliff I. Davidson. 2018. Pluvial flood risk and opportunities for resilience. WIREs Water 5, 6 (2018), e1302. arXiv:https://wires.onlinelibrary.wiley.com/doi/pdf/10.1002/wat2.1302 doi:10.1002/wat2.1302 [59] Cédric Roussel and Klaus Böhm. 2025. Introducing Geo-Glocal Explainable Artificial Intelligence. IEEE Access 13 (2025), 30952–30964. doi:10.1109/ACCESS. 2025.3541781 [60] Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 1, 5 (May 2019), 206–215. doi:10.1038/s42256-019-0048-x [61] Aishwarya Sarkar, Autrin Hakimi, Xiaoqiong Chen, Hai Huang, Chaoqun Lu, Ibrahim Demir, and Ali Jannesari. 2025. HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction. In Proceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems (The Graduate Hotel Minneapolis, Minneapolis, MN, USA) (SIGSPATIAL ’25). Association for Computing Machinery, New York, NY, USA, 1019–1030. doi:10.1145/3748636.3764172 [62] D. W. Shin, Steven Cocke, and Baek-Min Kim. 2022. A Systematic Revision of the NFIP Claims Hazard Data in Florida for Flood Risk Assessment. Applied Sciences 12, 7 (Jan. 2022), 3537. doi:10.3390/app12073537 [63] Xinyi Shu, Zongxue Xu, Silong Zhang, Chenlei Ye, and Lei Yu. 2026. Assessing Pluvial Flooding Risk in Urban Areas with High Spatial Heterogeneity Using a Fused Physically-Based and Data-Driven Framework. International Journal of Disaster Risk Science (May 2026). doi:10.1007/s13753-026-00728-8 [64] Sanah Suri and Maike Sonnewald. 2026. Trusting machine learning with physics: A fidelity verification framework for complex systems. ESS Open Archive 2026, 0120 (2026). doi:10.22541/essoar.176894678.89831689/v1 [65] Daniela Szwarcman, Sujit Roy, Paolo Fraccaro, Þorsteinn Elí Gíslason, Benedikt Blumenstiel, Rinki Ghosal, Pedro Henrique de Oliveira, Joao Lucas de Sousa Almeida, Rocco Sedona, Yanghui Kang, Srija Chakraborty, Sizhe Wang, Carlos Gomes, Ankur Kumar, Vishal Gaur, Myscon Truong, Denys Godwin, Sam Khallaghi, Hyunho Lee, Chia-Yu Hsu, Ata Akbari Asanjan, Besart Mujeci, Disha Shidham, Rufai Omowunmi Balogun, Venkatesh Kolluru, Trevor Keenan, Paulo Arevalo, Wenwen Li, Hamed Alemohammad, Pontus Olofsson, Timothy Mayer, Christopher Hain, Robert Kennedy, Bianca Zadrozny, David Bell, Gabriele Cavallaro, Campbell Watson, Manil Maskey, Rahul Ramachandran, and Juan Bernabe Moreno. 2026. Prithvi-EO-2.0: A Versatile Multitemporal Foundation Model for Earth Observation Applications. IEEE Transactions on Geoscience and Remote Sensing 64 (2026), 1–20. doi:10.1109/TGRS.2025.3642610 [66] Mehdi Taghizadeh, Zanko Zandsalimi, Mohammad Amin Nabian, Majid Shafiee-Jood, and Negin Alemazkoor. 2025. Interpretable physicsinformed graph neural networks for flood forecasting. ComputerAided Civil and Infrastructure Engineering 40, 18 (2025), 2629– 2649. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/mice.13484 doi:10.1111/mice.13484 [67] Vinh Ngoc Tran, Taeho Kim, Donghui Xu, Hoang Tran, Manh-Hung Le, ThanhNhan-Duc Tran, Jongho Kim, Trung Duc Tran, Daniel B. Wright, Pedro Restrepo, and Valeriy Y. Ivanov. 2025. AI Improves the Accuracy, Reliability, and Economic Value of Continental-Scale Flood Predictions. AGU Advances 6, 3 (June 2025), e2025AV001678. doi:10.1029/2025AV001678

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

[68] Pankaj Upreti and CSP Ojha. 2021. Comparison of antecedent precipitation based rainfall-runoff models. Water Supply 21, 5 (2021), 2122–2138. [69] U.S. Army Corps of Engineers. 2022. National Structure Inventory (NSI) Base Data. https://www.hec.usace.army.mil/confluence/nsi. [70] U.S. Geological Survey (USGS). 2024. Annual NLCD Collection 1 Science Products. doi:10.5066/P94UXNTS [71] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 6000–6010. [72] Markus Weiler, Julia Krumm, Ingo Haag, Hannes Leistert, Max Schmit, Andreas Steinbrich, and Andreas Hänsler. 2025. The Pluvial Flood Index (PFI): a new instrument for evaluating flash flood hazards and facilitating real-time warning. doi:10.5194/egusphere-2025-1519 [73] Oliver E. J. Wing, Paul D. Bates, Niall D. Quinn, James T. S. Savage, Peter F. Uhe, Anthony Cooper, Thomas P. Collings, Nans Addor, Natalie S. Lord, Simbi Hatchard, Jannis M. Hoch, Joe Bates, Izzy Probyn, Sam Himsworth, Josué Rodríguez González, Malcolm P. Brine, Hamish Wilkinson, Christopher C. Sampson, Andrew M. Smith, Jeffrey C. Neal, and Ivan D. Haigh. 2024. A 30 m Global Flood Inundation Model for Any Climate Scenario. Water Resources Research 60, 8 (2024), e2023WR036460. doi:10.1029/2023WR036460 _eprint: https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1029/2023WR036460. [74] Oliver E. J. Wing, Nicholas Pinter, Paul D. Bates, and Carolyn Kousky. 2020. New insights into US flood vulnerability revealed from flood insurance big data. Nature Communications 11, 1 (March 2020), 1444. doi:10.1038/s41467-020-15264-2 [75] Sean A. Woznicki, Jeremy Baynes, Stephanie Panlasigui, Megan Mehaffey, and Anne Neale. 2019. Development of a spatially complete floodplain map of the conterminous United States using random forest. Science of The Total Environment 647 (Jan. 2019), 942–953. doi:10.1016/j.scitotenv.2018.07.353 [76] Qing Yang, Xinyi Shen, Feifei Yang, Emmanouil N. Anagnostou, Kang He, Chongxun Mo, Hojjat Seyyedi, Albert J. Kettner, and Qingyuan Zhang. 2022. Predicting Flood Property Insurance Claims over CONUS, Fusing Big Earth Observation Data. Bulletin of the American Meteorological Society 103, 3 (March 2022), E791–E809. doi:10.1175/BAMS-D-21-0082.1 [77] Stefano Zanardo and Jose Luis Salinas. 2022. An introduction to flood modeling for catastrophe risk management. WIREs Water 9, 1 (2022), e1568. arXiv:https://wires.onlinelibrary.wiley.com/doi/pdf/10.1002/wat2.1568 doi:10.1002/wat2.1568 [78] Lin Zhang, Huapeng Qin, Junqi Mao, Xiaoyan Cao, and Guangtao Fu. 2023. High temporal resolution urban flood prediction using attention-based LSTM models. Journal of Hydrology 620 (May 2023), 129499. doi:10.1016/j.jhydrol.2023.129499 [79] Xiao Xiang Zhu, Zhitong Xiong, Yi Wang, Adam J. Stewart, Konrad Heidler, Yuanyuan Wang, Zhenghang Yuan, Thomas Dujardin, Qingsong Xu, and Yilei Shi. 2026. On the foundations of Earth foundation models. Communications Earth & Environment 7, 1 (Jan. 2026), 103. doi:10.1038/s43247-025-03127-x

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

A

Data Sources

This appendix gives per-source descriptions of the input modalities summarized in §4.2, organized by the risk component each contributes to.

A.1

Hazard: Hydrometeorology

NOAA Analysis of Record for Calibration (AORC). High quality record of precipitation with sufficient temporal and spatial resolution is critical to account for flood damages. Past works have employed a variety of precipitation records including NASA’s IMERG [31] or ERA5 [28] reanalysis data; however, these datasets are insufficient for the daily, ∼1km scale of our analysis. In this work, we use the NOAA Analysis of Record for Calibration (AORC) dataset [21], which provides a high quality, high resolution (hourly, ∼800m) precipitation record for the Conterminous United States. We use the Hourly Accumulated Precipitation record from this archive. NOAA National Water Model (NWM) Retrospective. Precipitation signal alone is often insufficient to model pluvial flood occurrence and intensity, as past works have shown that pluvial flooding is fundamentally a hydrological problem [34]. Thus to supplement the NOAA AORC precipitation records, we turn to NOAA’s National Water Model (NWM) Retrospective dataset v3.0 [18] available from Feb. 1979 to Jan. 2023. Amongst its various outputs, we use the Land Surface Model component, which provides 3-hourly records of physical land surface and hydrologic states at ∼1 km resolution. We use two variables from the NWM Retrospective dataset: Volumetric Soil Moisture and Snow Water Equivalent.

A.2

Hazard: Terrain

Topographic Wetness Index (TWI). Topographic Wetness Index quantifies the steady-state propensity of a location to accumulate surface water from its upslope contributing area, defined as TWI = ln(𝑎/tan 𝛽), where 𝑎 is the upslope contributing area per unit contour length and 𝛽 is the local slope. We use the CONUSwide 30m product from [29]. Height Above Nearest Drainage (HAND). Height Above Nearest Drainage measures the vertical elevation of each cell above its hydrologically connected drainage network. Low-HAND cells lie nearest the drainage corridors that inundate first when local drainage capacity is exceeded, making HAND a direct proxy for pluvial flood susceptibility. We use the Copernicus DEM HAND dataset at 30 m resolution [20]. Global Curve Number (GCN). The SCS Curve Number parameterizes surface runoff potential as a joint function of soil hydrologic group and land cover, with values ranging from ∼30 (high infiltration on sandy or forested terrain) to 100 (fully impervious surfaces). Higher CN indicates a larger fraction of incident precipitation converted to surface runoff, the proximate hydrologic driver of pluvial flooding. CN values from [33] are derived from three different antecedent runoff conditions (ARC): dry (ARC I), average (ARC II), and wet (ARC III) soil moisture states, which we include as separate features. We use the three Global Curve Numbers as terrain hazard descriptors.

Kawakami et al.

NLCD Land Cover. Land cover plays an important role in pluvial flood risk with impervious surfaces generating more runoff and thus more pluvial flood risk. To account for impervious surfaces, we include the National Land Cover Database (NLCD) 2016 product [70], which provides 30m land cover classifications across the Conterminous United States. We use the Fractional Impervious Surface classification as a surface descriptor. POLARIS Soil Properties. Soil plays a critical role in modeling pluvial flood events [57], as it controls the partitioning of precipitation into infiltration and runoff. To account for the varied soil properties across the Conterminous United States, we include the POLARIS dataset [15], which provides high-resolution (30m) estimates of soil properties such as texture, hydraulic conductivity, and water retention characteristics. From the dataset, we identify three variables of interest: 𝐾𝑠𝑎𝑡 , the saturated hydraulic conductivity (log10 (𝑐𝑚/ℎ𝑟 )); clay, the clay percentage (%); and 𝜃 𝑠 , the saturated soil water content (𝑚 3 /𝑚 3 ). All values are derived from the 0–5 cm soil layer.

A.3

Exposure and Vulnerability

FIMA NFIP Policies. The FIMA NFIP policies dataset provides spatially aggregated counts and total coverage of active flood insurance policies across the Conterminous United States. We use policy density and total insured value as exposure proxies; regions with denser NFIP coverage contain more structures with known flood vulnerability and a larger insured asset base at risk. These are derived from the OpenFEMA FIMA Redacted Policies Dataset v2 [25] and undergo the same spatial uncertainty correction as the claims data (§3.1.2), yielding a polygon-level count and total coverage for each cell in our analysis grid. We derive # Active NFIP Policy from this archive. National Structure Inventory (NSI). The U.S. Army Corps of Engineers National Structure Inventory [69] provides a point-level inventory of structures across CONUS, with per-structure attributes including replacement value, occupancy type, square footage, firstfloor elevation, and foundation type. From this archive we derive and rasterize using the same methodology as in §3.1.2 the following variables: Total Count of Structures, % Structs with Basements, % Struct. with Slab Foundation, % Mobile Homes, % Structs in Special Flood Hazard Area (SFHA), and Average, Min, Standard Deviation First Floor Elevation.

A.4

Google AlphaEarth Foundation (AEF)

The AlphaEarth Foundation model provides a 64-dimensional learned embedding of Earth’s surface at 10 m resolution, trained on multisource Earth observation imagery and other environmental datasets [13]. AEF embeddings implicitly encode land cover, impervious surface fraction, urbanization patterns, vegetation, and built-environment morphology, all factors that modulate runoff generation and drainage capacity, providing a complementary, dense representation alongside the hand-engineered terrain and exposure features above.

A.5

Normalization

To stabilize training, we apply transformations to our raw feature values. For those that exhibit high right skew, HAND, TWI, #

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

Active NFIP Policies, # NSI Structures, we apply a ln(𝑥 + 1) transformation followed by a z-score normalization. For NSI Foundation height features, we apply a z-score normalization without the log transform, as these features are not as heavily skewed. For the rest we apply a min-max normalization to scale values to the [0,1] range.

B

Date Shift Distribution

We describe our temporal correction procedure in §3.1.1. Here, we report the distribution of date shifts Δ = 𝑑 ′ − 𝑑 applied to the original NFIP claim dates 𝑑 to obtain the corrected dates 𝑑 ′ used in our study region. As shown in Table 6, the vast majority of claims (81.1%) had no date shift, while 7% has shifts of ±1 day and 6% were discarded. Δ (days)

−3

−2

−1

0

+1

+2

+3

%

2.0

1.9

5.0

81.1

2.0

1.4

0.6

6.0

Table 6: Distribution of date shifts Δ = 𝑑 ′ − 𝑑 from the temporal correction. (See §3.2)

C

All Held-Out Cell Modulator Dynamics

For completeness, Figures 8, 9 and 10 show, respectively, the basemap, the AEF clustering, and the learned modulator clustering for every held-out cell. Panel positions match across the three figures.

D

Temporal Generalization

Under a strict temporal split (train 2017–2020, test 2021–2022, single split), DELUGE retains its lead on both PR-AUC and $-Weighted PR-AUC (Table 7). PR-AUC degrades only modestly from the spatial split, with DELUGE the least affected at a ∼3% drop and the tree baselines down ∼7 to 10%, preserving DELUGE’s lead and the baseline ordering. $-Weighted PR-AUC drops sharply for every model (roughly ∼38 to 41% from the spatial split), but DELUGE retains its lead and the ordering across baselines is preserved. The Model DELUGE (Ours) XGBoost LightGBM

PR-AUC

Δ

$-W PR-AUC

Δ

0.237 0.197 0.212

– −17% −11%

0.370 0.303 0.329

– −18% −11%

Table 7: Temporal split (train 2017–2020, test 2021–2022, single split). DELUGE retains its lead on both metrics. divergence between metrics is structural. PR-AUC averages over all positives whereas $-Weighted PR-AUC concentrates on the few costliest claims, so models are judged primarily on how well they rank the tail. Training is dominated by Hurricane Harvey in 2017 ($5.8B, more than the other three training years combined at $2.0B), a clear signal that any reasonable model can learn to flag. Test years ($1.4B in 2021, $1.9B in 2022) lack a comparably obvious extreme, which hampers the $-Weighted PR-AUC of all models.

E

Performance Comparison by Location Characteristics

To understand whether DELUGE’s headline performance is uniform across our study cells or concentrated in particular locations, we stratify the held-out cells by their underlying landscape and

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Cluster C1: NE developed C2: Coastal wetland C3: Inland developed C4: Dense developed C5: SE mixed rural-developed C6: Inland less-developed

Prev. (%)

PR-AUC

Lift

0.139 0.166 0.065 0.610 0.249 0.222

0.184 0.138 0.080 0.255 0.322 0.238

132× 83× 124× 42× 129× 107×

Table 8: DELUGE performance stratified by AEF cluster. Heldout cells grouped by K-means (𝐾 = 6) on the center-pixel AlphaEarth embedding. Prev. is per-pixel-day prevalence and lift is PR-AUC / prevalence, the multiplicative gain over a random predictor. report PR-AUC within each stratum. We characterize each held-out pixel by its AlphaEarth Foundation embedding [13] and cluster the resulting per-pixel vectors with K-means and a silhouette sweep selects 𝐾 = 6. We then report DELUGE’s PR-AUC within each AEF cluster (Table 8), along with the lift over prevalence (PR-AUC / prevalence), the multiplicative gain over a random predictor. Inspecting the clusters’ geographic membership, we label them as seen in Table 8. DELUGE’s raw PR-AUC varies substantially across these strata. Lift over prevalence largely controls for this, and across most clusters the gain over chance is comparable. Of note, however, is the poor relative lift of the dense developed cluster (C4, 42×). Despite the highest raw PR-AUC of any builtenvironment stratum, DELUGE’s marginal contribution over a baserate predictor is the smallest here. We attribute this to local, subkilometer flood dynamics in dense urban areas, that operate below the ∼1 km resolution at which DELUGE operates. This matches the localization failure mode to be discussed in §5.4 and points to spatial resolution as a natural lever for future work in this regime. We place the full cluster membership and geographic distribution in Appendix. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Figure 8: Basemaps of all 30 held-out cells, paired with the modulator kernels in Figure 10

Kawakami et al.

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction

Cluster 1

Cluster 2

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Cluster 3

Cluster 4

Cluster 5

Cluster 6

Figure 9: Google AlphaEarth Clustering results for all 30 held-out cells. See §E.

Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Figure 10: Learned modulator clustering for all 30 held-out cells. See §5.3.

Kawakami et al.

Record · ID 381739 · SHA-256 f16b9e0b0b182535
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.