Chronosphere: Space-Time Tessellation of Local Climate Experts Daniel Cher1 , Eric Xing1 , Kexing Li1 , Brian Wei1 , Isaac Corley2 , Nathan Jacobs1 1
arXiv:2609.21872v1 [cs.CV] 18 Sep 2026
Washington University in St. Louis 2 Taylor Geospatial Institute {cher, e.xing, ariana.l, b.j.wei, jacobsn}@wustl.edu, [email protected]
JAN
APR
NOV
JUL
Figure 1: Chronosphere is a single learned spatio-temporal representation of Earth’s climate. Each globe shows the frozen embedding queried at a different point in the year (Jan → Dec clockwise from the top), projected to RGB over land. The embedding shifts smoothly from month to month, so one representation captures both where a place sits and how it changes through the year. Abstract We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail everywhere. Global bases such as spherical harmonics spread capacity uniformly across space and time. Localized bases resolve only predefined regions. Learned tessellations adapt, but are inefficient at representing higher frequencies. Chronosphere unifies these approaches, pairing an adaptive tessellation of learnable sites on the spacetime torus S 2 × S 1 with a shared bank of local basis functions. Both where capacity is placed and how much detail each region carries adapt to the data, across space and time. Trained to reconstruct climatology, Chronosphere matches or leads stateof-the-art location encoders across spatial and temporal tasks, with the largest gains under spatial and temporal transfer.
Introduction Location encoders have become a popular tool in geospatial modeling, turning a coordinate into a reusable representa-
tion of the environment. They have proven useful across tasks such as species distribution modeling (Cole et al. 2023; Mac Aodha, Cole, and Perona 2019), crop-yield estimation (Tseng et al. 2025), air-quality forecasting (Karimzadeh, Wang, and Crooks 2025), canopy-height mapping (Lang et al. 2023), and satellite image synthesis (Sastry et al. 2024; Cher et al. 2026b; Wei et al. 2026). Much of this work has focused on well-sampled settings, where a probe fit on the representation interpolates among nearby labels. But these representations matter most where local labels are scarce, when a task must generalize to a distant region or an unseen season (Cai and Balestriero 2025). That need is sharpening as climate conditions migrate across the landscape (Loarie et al. 2009; Milly et al. 2008) and the timing of the seasons shifts with them (Burrows et al. 2011; Williams and Jackson 2007), carrying the environment away from where any model was trained. A useful representation must therefore generalize across space and time, which depends on capturing how the environment itself varies. Environmental fields vary in frequency across both space and season (Rolf et al. 2024). Temperature varies little across
Spacetime Torus (Learned Spatiotemporal Sites: 𝑠!,# ) Learnable Spherical Voronoi (Site embeddings: 𝑒!,# )
Mixture of Fields 𝑧(𝑥, 𝑡)
∑ Displacement 𝒅
…
…
𝑥 𝑠!,#
× Temporal Coefficients ∝!,$ (𝑡)
𝑑 = 𝑥 − 𝑠!,# Expert Bank 𝑓$ (𝑑)
Localized Field Experts
Figure 2: Chronosphere. A query, a location x ∈ S 2 and within-year time t, is softly assigned to learnable sites on the spacetime torus S 2 × S 1 . At each site a shared bank of sine experts fj is mixed by per-site coefficients αk,j (t) and added to the aggregated site embedding to form the location vector z(x, t), which the decoder reconstructs into climate fields. the Amazon basin but changes sharply over the neighboring Andes. Likewise, the Sahel remains nearly uniform for much of the year until the onset of the monsoon produces a strong moisture gradient within a matter of weeks. A representation should therefore concentrate resolution where the field is sharp and stay coarse where it is smooth, rather than fixing the level of detail across the globe and year. Existing encoders fall short of this in different ways. Global bases such as spherical harmonics cover the sphere but spread capacity uniformly (Rußwurm et al. 2024), localized bases such as Slepian functions concentrate resolution only within predefined regions (Rao et al. 2026), and learned tessellations let sites adapt to the data but gain finer structure only by adding more sites (Cher et al. 2026a), never enriching each one. Temporal signals enter as uniform global conditioning applied identically at every location (Dollinger et al. 2025; Mickisch et al. 2025). These approaches have been pursued in isolation, but their strengths are complementary. Combining the parameter efficiency of a shared basis with the adaptivity of a learned tessellation, without predefined regions, gives a representation that generalizes where labels are sparse. We introduce Chronosphere (Fig. 1), a spatio-temporal neural field built on this idea (Fig. 2). Its learnable sites live on the spacetime torus S 2 × S 1 , each carrying a spatial and temporal position, and queries are encoded as soft Voronoi weightings over joint spatio-temporal distance. Each site draws on a shared bank of local basis-function experts spanning a range of frequencies, mixing them differently through the year to set how finely its region is resolved and how that detail shifts with the seasons. Trained to reconstruct climatology, the frozen embedding learns fine-grained spatial variation across the globe (Fig. 2) and seasonal variation
that matches the geography of the annual cycle (Fig. 5). Because these representations matter most where labels are scarce, we evaluate under both interpolation and transfer, using spatial block and temporal holdouts. Chronosphere matches or leads existing geospatial and spatiotemporal encoders across a range of spatial and temporal tasks, with its largest advantage on interpolation of temporal tasks and under spatial or temporal transfer, where local autocorrelation is minimal. The main contributions of this work are: • Chronosphere, a spatio-temporal neural field that unifies a learnable Voronoi tessellation with a bank of local basis functions, adapting capacity and detail across space and time. • An evaluation across a suite of spatial and temporal tasks, where Chronosphere matches or beats the best encoders in-distribution and leads them under generalization to unseen regions and seasons. • Qualitative and quantitative evidence that the learned representation matches real climate structure, from finegrained spatial fields to the phase and geography of the yearly cycle.
Related Work Spatio-Temporal Representation Learning Location encoders differ most in the signal that shapes their embeddings (Mai et al. 2022). Many align coordinates with street-level (Vivanco Cepeda, Nayak, and Shah 2023; Yin et al. 2019), satellite (Klemmer et al. 2025; Cher et al. 2026a), and citizen-science (Mai et al. 2023a) imagery through contrastive learning, while SINR (Cole et al. 2023) instead trains
on species occurrence, and RANGE (Dhakal et al. 2025) adds retrieval augmentation to recover the fine spatial detail contrastive pretraining tends to lose. These encoders yield useful representations of places, but each maps a location to a single spatial embedding and none capture how that location changes through the year. Models that do incorporate time typically add it as a global signal. CLIMPLICIT (Dollinger et al. 2025) pretrains a location encoder to regress monthly CHELSA climatology, so that its embedding is organized by climate and carries a seasonal signal. That signal enters as a single global function conditioned on the month, rather than as structure that varies from place to place. Joint space-time encoders (Mickisch et al. 2025; Shatwell et al. 2025) learn location and time together in a single network, but still fix how finely each place is resolved rather than adapting it across space and season. Chronosphere shares CLIMPLICIT’s climate-field objective but builds season into the site geometry itself, giving the representation its own spatio-temporal structure.
Adaptive Capacity Allocation in Neural Fields Most encoders allocate representational capacity through a global basis over the whole sphere, whether multi-scale sinusoids (Mai et al. 2020, 2023b) or spherical harmonics with a SIREN (Rußwurm et al. 2024), spreading it uniformly. Slepian bases (Rao et al. 2026) instead concentrate resolution within regions fixed in advance. Learnable partitions make this allocation adaptive. Spherical Voronoi tessellations and related decompositions (Di Sario et al. 2026; Rebain et al. 2021; Cher et al. 2026a) learn where sites are placed so that capacity follows the data, yet each site holds a single smooth vector with fixed within-region detail. Local implicit fields (Sitzmann et al. 2020; Tancik et al. 2020) supply that detail, but as one global function with no partition to place it. Each gives either an adaptive partition or rich within-region detail, never both. Chronosphere unifies them, pairing a learnable partition with a bank of local basis-function experts so that both where capacity sits and how finely each region resolves adapt to the data, jointly across space and season.
Data Chronosphere uses two complementary climate corpora, monthly CHELSA climatology (∼1 km, sharp spatial structure) and daily ERA5-Land reanalysis (∼9 km, within-year and across-year dynamics). The monthly variant trains on CHELSA alone, while the daily variant trains jointly on both. Downstream evaluation datasets are described in the Experimental Setup section (full provenance tables are deferred to the appendix). Monthly climatology (CHELSA). We train the monthly CHELSA branch on v2.1 monthly climatologies (WMO 1981–2010 normal) at ∼1 km over land (Karger et al. 2017). Because this is a climatological normal, the corpus emphasizes spatial structure and monthly seasonality without yearto-year variation. Following CLIMPLICIT (Dollinger et al. 2025), the model regresses eleven normalized variables spanning temperature, precipitation, radiation, wind, humidity,
and derived moisture indices, listed in full in the appendix (Training Data Variables). Daily reanalysis (ERA5-Land). The ERA5 branch uses daily sequences (1970–2024) for within-year and yearresolved modeling (Muñoz-Sabater et al. 2021). It covers the same land climate variables as CHELSA, minus cloud cover, but at a coarser ∼9 km native resolution that limits fine-scale spatial detail. Climate indices (inter-annual drivers). Daily reanalysis captures seasonal structure, but year-to-year variation (slow forced trends and shorter-lived anomalies) requires an explicit inter-annual signal. A free per-year embedding can memorize in-sample years but does not define years out of distribution. We condition the daily model on eight monthly climate indices (ENSO, NAO, PDO, AMO, global mean temperature, log CO2 , stratospheric aerosol, and sunspot activity): large-scale drivers whose reconstructions extend to the mid-19th century, so the same conditioning can apply in training, backcast, and forecast once index values are supplied. The indices are inputs, not reconstruction targets. We describe how they modulate the predictions in the Hybrid inter-annual head subsection.
Experimental Setup Downstream tasks. We evaluate frozen encoders on downstream tasks chosen for a strong climate-driven signal, grouped by the temporal structure they require (Table 2). Spatial niche tasks (WHERE) ask what kind of place a coordinate represents: biome, ecoregion, and elevation from the RANGE benchmark (Dhakal et al. 2025), plant functional traits (Lusk et al. 2026), and canopy height (Lang et al. 2023). Within-year tasks (WHEN) probe the seasonal cycle via MODIS NDVI (Didan 2015), snow cover (Hall and Riggs 2015), river discharge (Kratzert et al. 2023), GFED burned area (van der Werf et al. 2017), and USA-NPN phenology (USA National Phenology Network 2012; Rosemartin et al. 2018). Table 2: Climate-driven downstream tasks (▼ WHERE, ⟳ WHEN). ▼ ▼ ▼ ▼ ▼ ⟳ ⟳ ⟳ ⟳ ⟳
Task
Target
Resolution
Biome Ecoregion Elevation Plant traits Canopy height MODIS NDVI MODIS snow Discharge GFED fire USA-NPN
Biome label Ecoregion label Elevation SLA, LDMC, LNC, LPC Canopy height (m) Greenness Snow cover Streamflow Burned area Phenophases
Static Static Static Static Static 16-day 8-day Daily 8-day Visit date
Evaluation protocol. For temporal tasks we evaluate frozen encoders at each task’s native resolution. We use the embedding at the matching day, month, or year and never pool across time. Spatial niche tasks assign one feature vector to each location. We run the encoder at twelve monthly or at
Table 1: Downstream performance (frozen encoders). Each cell is iid / extrapolation; bold = best, italic = second best, ranked separately within each half. Scores are means over three seeds. Chronosphere’s per-task standard deviation stays under 0.01 on every iid split. Under extrapolation it stays below ∼0.03 on most tasks, rising to 0.04–0.06 on the few most sensitive to the split (plant traits, snow, elevation, streamflow). Spatial (iid / 10◦ geo-extrap.) Acc. ↑
↑
Within-year (iid / window-extrap.) R2 ↑
AUC ↑
AP ↑
Model
Bio
Eco
Elev
Traits
Canopy
NDVI
Snow
Streamflow
Fire
USA-NPN
SINR CSP-iNat GeoCLIP SatCLIP TTE
0.67/0.55 0.59/0.46 0.60/0.41 0.69/0.54 0.77/0.59
0.48/0.12 0.54/0.08 0.60/0.09 0.69/0.20 0.67/0.28
0.57/−0.05 0.34/0.13 0.39/−0.06 0.66/0.32 0.84/0.65
0.59/0.12 0.42/0.20 0.37/0.15 0.50/0.31 0.67/0.41
6.1/7.3 7.0/8.0 8.2/9.6 5.9/6.3 5.1/5.6
0.68/0.59 0.56/0.47 0.41/0.40 0.69/0.60 0.76/0.66
0.48/−0.30 0.44/−0.33 0.39/−0.31 0.45/−0.29 0.50/−0.28
0.08/−0.38 0.08/−0.37 0.05/−0.36 0.10/−0.38 0.10/−0.38
0.81/0.79 0.77/0.76 0.75/0.75 0.82/0.80 0.84/0.82
0.11/0.09 0.10/0.09 0.11/0.10 0.11/0.09 0.11/0.10
GT-Loc STE CLIMPLICIT
0.72/0.59 0.83/0.58 0.81/0.58
0.71/0.14 0.75/0.32 0.79/0.28
0.62/0.31 0.88/0.69 0.89/0.75
0.51/0.39 0.62/0.17 0.57/0.30
5.8/6.3 5.0/5.9 5.1/5.6
0.71/0.67 0.79/0.77 0.77/0.73
0.57/0.03 0.78/0.68 0.81/0.70
0.21/−0.13 0.50/0.21 0.42/0.12
0.82/0.81 0.84/0.84 0.81/0.78
0.20/0.11 0.23/0.11 0.22/0.09
Chronosphere (monthly) Chronosphere (daily)
0.83/0.65 0.81/0.66
0.78/0.31 0.76/0.28
0.94/0.87 0.92/0.78
0.67/0.47 0.66/0.47
4.5/5.1 4.7/5.0
0.87/0.84 0.87/0.86
0.85/0.76 0.89/0.85
0.53/0.31 0.58/0.40
0.88/0.87 0.88/0.87
0.23/0.12 0.25/0.12
R
MODIS NDVI (greenness)
RMSE ↓
Burned area
MODIS snow cover within-year R 2
0.7 0.6 0.5
burned/no-burned AUC
0.8
0.8
within-year R 2
2
0.6 0.4 0.2 0.0 0.2
0 10 20
40
60
seasonal-phase gap (days)
90
Chronosphere (daily)
0 10 20
40
60
seasonal-phase gap (days)
Chronosphere (monthly)
STE
90
0.86 0.84 0.82 0.80 0.78 0.76 0.74 0.72
0 10 20
40
60
seasonal-phase gap (days)
GT-Loc
90
CLIMPLICIT
Figure 3: Temporal extrapolation across seasonal phase. Probe skill as the seasonal-phase gap δ between the training periods and the held-out window grows, from in-distribution (left) to a full season held out (right). Greenness and snow are scored by R2 and burned area by burned/no-burned ROC-AUC, with the dotted line marking chance. Time-aware encoders only. twenty-four day-of-year steps for temporal models, then concatenate the mean, standard deviation, minimum, and maximum so the linear probe captures local seasonality, while retaining a consistent feature dimension across models and tasks. We follow the previous testing protocols (Dhakal et al. 2025; Klemmer et al. 2025). Every task is scored with a cross-validated ridge probe, ridge classification for the categorical targets and ridge regression for the continuous ones, so all encoders are compared under one fixed protocol. The appendix (RANGE Linear Protocol) gives the full protocol. Our main table reports interpolation on three random 80/20 splits and spatial extrapolation on held-out 10◦ lat/lon tiles. We also report results at tile sizes of 5◦ → 80◦ to trace how accuracy changes with geographic separation. A random split places test points inside the spatial and temporal autocorrelation range of the training data, so interpolation scores largely reflect memorized local context and can overstate how far an encoder transfers (Roberts et al. 2017; Kattenborn et al.
2022; Ploton et al. 2020). Blocked holdouts remove that leakage. Separating whole tiles in space and contiguous windows in season measures the representation’s value where autocorrelation gives nothing to lean on, the regime that matters as the climate moves into unseen states (Valavi et al. 2019). Within-year evaluations hold out contiguous blocks of the seasonal cycle.
Methodology Chronosphere maps a unit-sphere coordinate x ∈ S 2 and a within-year time t to a site embedding for downstream probing (Fig. 2). It lets the representation adapt capacity and detail where necessary across space and time. We train two variants of the same architecture. The monthly model reconstructs CHELSA climatologies alone at ∼1 km (P =12). The daily model jointly reconstructs CHELSA and ERA5-Land (P =365). Our flagship model uses K=8192 sites and E=48 experts over six frequency bands.
iid extrap.
Chrono (monthly) Chrono (daily)
1.0 TTE CLIMPLICIT STE
Spatial skill
Spatiotemporal sites. Each site behaves like a weather station pinned to both a place and a time of year, and a query, itself a place x and a day t, is described by a soft blend of the stations nearest it in both. We place K learnable sites, each with a direction sk ∈ S 2 , a day-of-year tk , a spatial temperature τk , a temporal weight βk , and an embedding ek ∈ RD , so each site is a point on the spacetime torus S 2 × S 1 . A single softmax assigns the query to sites and averages their embeddings into a location vector z0 (x, t), wk (x, t) = softmaxk τk (x · sk ) − βk δk (t) , | {z } | {z } spatial closeness temporal distance (1) X z0 (x, t) = wk (x, t) ek ,
SatCLIP
0.5
SINR CSP-iNat
params
k∈K
where δk (t) = 12 1 − cos(2π(t − tk )/P ) measures cyclic distance in the year and is zero when the query is the site’s day tk . Temporal resolution then emerges as spatial resolution does. Spatial detail comes from many sites at different directions competing for a query, and fine seasonal detail from sites that overlap in space but sit at staggered days of the year. At one location a summer-anchored and a winter-anchored site hold different embeddings, and a query’s weight crossfades between them as the year turns, continuously across the year boundary. This adds a temporal axis to the soft Voronoi tessellation (Cher et al. 2026a). We tie the seasonal weight to the spatial scale as βk = τk eγk and initialize it small with the tk spread across the year, so space dominates at first and the assignment reduces to the purely spatial encoder as βk → 0, adding seasonal structure only where the data supports it. Two simpler ways to add time, passing the day-of-year only to the decoder or evolving site embeddings through a low-rank seasonal flow, are compared in Table 3. Local Field Experts. A tessellation alone gives each site a single vector, so sharper structure would require adding sites, which quickly becomes computationally expensive. We instead add local texture as a continuous per-site field from a shared bank of sine experts (Fig. 2). Each expert fj is evaluated on a normalized displacement from the site that puts every site on a common unit scale, √ (2) dk (x) = τk (x − sk ). A site then forms its local field as a weighted sum of the experts’ outputs fj dk (x) (Eq. 2), with its own coefficients αk,j (t) and a readout Wj . These per-site fields are blended by the same soft assignment wk over a small fixed neighborhood of the query’s m nearest sites Tm (x), r(x, t) =
X k∈Tm (x)
wk (x, t)
E X
αk,j (t) Wj fj dk (x) , (3)
GT-Loc
2M 6M 12M
GeoCLIP
0.0 0.0
0.5 Within-year skill
1.0
Figure 4: In-distribution versus out-of-distribution performance. Each model runs from a hollow marker at its indistribution skill to a filled marker under extrapolation, with marker area the parameter count. A short segment toward the top-right means the encoder generalizes across both axes, and a long segment toward the bottom-left means it collapses out of distribution. Each axis is the min–max-normalized mean across that half’s tasks (best model ≈1, worst ≈0). Table 3: Chronosphere ablations (within-year, iid / window-extrap.). Chronosphere (daily) is from Table 1. Spatial ablations are in the appendix (Architectural Ablations). setting
NDVI
Snow
Streamflow
Fire
NPN
Chronosphere (daily)
0.87/0.86 0.89/0.85
0.58/0.40
0.88/0.86 0.25/0.12
Time in assignment (vs. spacetime) cyclic 0.85/0.81 0.84/0.73 flow 0.83/0.81 0.83/0.76
0.57/0.37 0.53/0.27
0.88/0.86 0.25/0.12 0.87/0.85 0.23/0.12
Local field (vs. free mixture) none (tessellation only) 0.82/0.79 0.86/0.79 gated 0.87/0.85 0.88/0.84 top-k 0.87/0.84 0.87/0.82
0.55/0.31 0.56/0.34 0.55/0.31
0.85/0.81 0.23/0.10 0.87/0.86 0.24/0.11 0.87/0.86 0.24/0.11
Site count K (vs. 8192) K=1024 0.87/0.85 K=2048 0.86/0.85 K=4096 0.86/0.85 K=16384 0.87/0.86
0.86/0.79 0.86/0.82 0.86/0.83 0.89/0.85
0.55/0.31 0.56/0.32 0.57/0.33 0.57/0.40
0.86/0.85 0.87/0.86 0.87/0.86 0.87/0.86
Frequency ladder (vs. full bank) low-freq 0.86/0.85 0.88/0.84 mid-freq 0.87/0.86 0.89/0.85 high-freq 0.87/0.86 0.90/0.86
0.56/0.38 0.57/0.39 0.58/0.40
0.87/0.86 0.24/0.12 0.87/0.86 0.24/0.12 0.88/0.86 0.23/0.11
0.24/0.12 0.24/0.12 0.24/0.12 0.25/0.12
j=1
|
{z
per-site local field
}
and the location vector is z(x, t) = z0 (x, t) + r(x, t). The coefficients αk,j (t) vary through the year, so a site’s local detail shifts with the season, while the experts fj are fixed spatial shapes shared across all sites. The readout Wj is
zero-initialized, so the residual r starts at zero and the model begins as the plain substrate, earning within-region detail only where the data supports it. The flagship uses sine experts, though the basis is a plug-in choice, with alternatives studied in the appendix (Architectural Ablations).
Jan
Apr
Jul
NE Siberia
2.0 1.5
kz(t) − z̄k
2.5
1.0
S-Central Africa Seasonal amplitude ∥z(t) − z̄∥
Figure 5: The location vector carries seasonal structure that differs in phase and character across the globe. Left: seasonal amplitude of the monthly vector, ∥z(t) − z̄∥, the distance the representation travels from its annual mean, with boxes marking the two regions shown at right. Right: the vector (PCA → RGB, one shared basis) in January, April, and July for NE Siberia and S-Central Africa. The two regions peak in opposite seasons and move along different color axes, one temperature-driven and one precipitation-driven, so a single shared embedding captures both when each cycle turns and what drives it. Decoder and embedding. Because time is already in z(x, t), the climatology decoder is a two-block residual MLP that reads z alone. Under joint training, the daily variant feeds the shared substrate to two reconstruction branches: a CHELSA head for eleven monthly variables and an ERA5 head for ten daily variables. The monthly variant uses CHELSA alone. We probe using the decoder’s frozen 256-dimensional penultimate activation. Hybrid inter-annual head. To backcast and forecast rather than reconstruct climatology alone, we add a small inter-annual head that puts a year-to-year residual on the climatology, driven by a set of standardized climate indices. The residual is linear in the indices, so raising one index by a fixed amount shifts the reconstruction by that index’s learned spatial pattern. The full parameterization and a comparison of alternatives are in the appendix (Yearly Index Drivers). Training. The monthly variant minimizes mean squared error on CHELSA targets over randomly sampled land locations and months. The daily variant jointly minimizes MSE on ERA5 and CHELSA samples, plus an ℓ2 prior that keeps the inter-annual residual small. Full hyperparameters, optimizer groups, and ablation configurations are in the appendix (Training Details).
Results and Discussion Downstream performance Chronosphere matches or leads across the spatial and temporal tasks in Table 1, and its margin widens under spatial and temporal held-out splits. Figure 4 collapses both axes into one view, where Chronosphere sits top-right with the shortest drop from interpolation to extrapolation. Its two variants specialize by resolution. The monthly model trains on CHELSA at ∼1 km and leads the spatial tasks. The daily model also trains on ERA5-Land at ∼9 km, trading spatial
detail for sub-monthly resolution, and leads the within-year tasks. Spatial tasks. Chronosphere leads every baseline on the three continuous spatial targets under 10◦ holdout: elevation (0.87), plant traits (0.47), and canopy height (5.1 m). The appendix (Spatial Extrapolation Across Tile Size) traces this as a curve, with Chronosphere highest at every separation. On the categorical targets it ties the best on biome and trails only on ecoregion, where CLIMPLICIT edges the iid split (0.79 against 0.78) and STE the extrapolation (0.32 against 0.31), each by one point. STE (Mickisch et al. 2025) is the closest competitor, but Chronosphere is the most consistent across targets and the strongest across all experimental regimes. Within-year tasks. The ordering follows how each encoder represents time. Encoders without a temporal axis hold one vector per location, so their probes cannot vary through the year and turn negative on snow and streamflow under window extrapolation. The time-aware baselines recover both, but Chronosphere leads them on all five within-year targets under both splits. Notably, while CLIMPLICIT and STE train on the same data as the monthly model, Chronosphere reaches 0.85 and 0.40 on snow and streamflow extrapolation, while CLIMPLICIT (0.70 and 0.12) and STE (0.68 and 0.21) trail by large margins. This advantage grows on NDVI, snow, and streamflow with finer, daily training data. In addition, as seen in Fig. 3, this lead widens under seasonal transfer. As the held-out window moves further from any training season, every encoder loses skill, but Chronosphere degrades least and stays above all other encoders. Beyond these climatology probes, a further experiment uses the frozen embedding as a spatio-temporal prior. We add it as an extra feature to an encoder-free baseline that predicts from past observations, on year-resolved records of fire, streamflow, water storage, dengue, greenness, and snow. Every climate encoder lifts this baseline, and Chronosphere
Warming trend
El Niño (temperature)
El Niño (rainfall)
Figure 6: Index-perturbation maps. Each panel raises one climate index by +1σ (the rest held at their mean) and shows the resulting change in the daily model’s mid-January reconstruction over land. Red is warmer or wetter, blue cooler or drier. lifts it most. The appendix (Encoder as a Geo-Prior) gives the full setup and results.
Ablations Essential structure. Table 3 ablates the daily model on the within-year targets, and two architectural decisions prove necessary. Putting time in the assignment (Eq. 1) beats passing the day-of-year to the decoder alone or drifting site embeddings through a seasonal flow. This is most clearly seen on snow extrapolation (0.85 against 0.73 and 0.76). Removing the local field (Eq. 3) leaves the plain tessellation and the worst row on every target, snow extrapolation falling to 0.79. An adaptive partition alone does not resolve structure within each region, so the partition and the local field are both essential. Prior encoders supply an adaptive partition or rich local detail but not both, and Table 1 confirms a global function without a partition trails on nearly every task. Implementation variants. The minimal form of each ingredient suffices. Free coefficients give each site direct control of its frequency content, and more elaborate mixing does not help. Gating adds a bottleneck and top-k breaks the smooth blend across experts, both losing ground on streamflow extrapolation (0.34 and 0.31 against 0.40). Adding capacity changes little with site count saturating by K=8192. Low frequencies alone fall behind, with most of the signal recovered once higher frequencies are included.
Learned temporal structure Within-year. Sampling the location vector through the year and measuring how far it departs from its annual mean, ∥z(t) − z̄∥, maps where the year moves the representation (Fig. 5, left). The vector changes most, as expected, where seasonality is strongest, over boreal Siberia and North America, the Sahel, and the South-Asian monsoon, and much less in the perennially wet tropics. We visualize a few regions in more detail over the year, converting each 256-dimensional vector to a color through a shared PCA projection to RGB (Fig. 5, right). In NE Siberia the vector sits at a cold, dark-blue extreme in January and swings to a warm yellow by July, a temperature-driven cycle along one PCA axis. Below the equator in S-Central Africa it peaks in the opposite season, lush green in the summer of January and red as the dry winter sets in by July, moving along a different, precipitation-driven axis. At the same calendar
month the two regions sit at opposite points of their annual loops. We read this as evidence that the embedding encodes both when each place’s cycle turns and what drives it. Inter-annual. Perturbing single indices reproduces known climate patterns. Greenhouse forcing warms the poles most, and El Niño shifts rainfall over Southeast Asia (Fig. 6). We show quantitative comparisons in the appendix (Yearly Index Drivers).
Conclusion We introduced Chronosphere, a spatio-temporal neural field that learns geographic representations, adapting to the data both where it spends capacity and how finely it captures detail. Its learnable sites live on the spacetime torus S 2 × S 1 , and a per-site mixture of local basis-function experts resolves detail within each region. Trained to reconstruct climatology, its embedding matches or leads state-of-the-art location encoders across spatial and temporal probes, with the largest margins under spatial and seasonal transfer. Prior encoders adopt only one of these two ingredients; unifying them drives the gains, and our ablations and baselines show that neither suffices alone. Several directions follow naturally from here. One is to extend to other spatio-temporal modalities such as satellite or street-level imagery, letting the model allocate resolution depending upon the change in the modality. A more comprehensive model could integrate many of these modalities together. In a similar vein, conditioning on the downstream task could move the model from a generalist representation toward one that learns the frequencies a given task depends on. The climatological model itself would also sharpen with better data. Known regions with sparse observational coverage (deserts, high mountains) rely heavily on interpolation, so those areas may not be as accurate. In addition, the daily variant model is spatially coarser than the monthly model, sacrificing temporal resolution for spatial. Above all we are excited for its use in ecology. Chronosphere lets researchers characterize the environment of almost any place and time, and we hope it becomes a useful representation for the many ecological problems that require climatological context.
Acknowledgments This research used the TGI RAILs advanced compute and data resource, which is supported by the National Science Foundation (award OAC-2232860) and the Taylor Geospatial Institute. This work is also supported by the Ann W. and Spencer T. Olin-Chancellor’s Fellowship at Washington University in St. Louis and by the AI-ACCESS National Research Traineeship, funded by the National Science Foundation (award DGE-2244165).
References Burrows, M. T.; Schoeman, D. S.; Buckley, L. B.; Moore, P. J.; Poloczanska, E. S.; Brander, K. M.; Brown, C. J.; Bruno, J. F.; Duarte, C. M.; Halpern, B. S.; Holding, J.; Kappel, C. V.; Kiessling, W.; O’Connor, M. I.; Pandolfi, J. M.; Parmesan, C.; Schwing, F. B.; Sydeman, W. J.; and Richardson, A. J. 2011. The Pace of Shifting Climate in Marine and Terrestrial Ecosystems. Science, 334(6056): 652–655. Cai, D.; and Balestriero, R. 2025. No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth Data. In The Thirteenth International Conference on Learning Representations (ICLR). Cher, D.; Iqbal, H.; Xing, E.; Wei, B.; and Jacobs, N. 2026a. Tessellating the Earth: Learnable Spherical Voronoi Partitions for Location Encoding. In European Conference on Computer Vision (ECCV). Cher, D.; Wei, B.; Sastry, S.; and Jacobs, N. 2026b. VectorSynth: Fine-grained satellite image synthesis with structured semantics. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 7019–7029. Cole, E.; Van Horn, G.; Lange, C.; Shepard, A.; Leary, P.; Perona, P.; Loarie, S.; and Mac Aodha, O. 2023. Spatial implicit neural representations for global-scale species mapping. In International conference on machine learning, 6320–6342. PMLR. Dhakal, A.; Sastry, S.; Khanal, S.; Ahmad, A.; Xing, E.; and Jacobs, N. 2025. RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings. In Proceedings of the Computer Vision and Pattern Recognition Conference, 24680–24689. Di Sario, F.; Rebain, D.; Verbin, D.; Grangetto, M.; and Tagliasacchi, A. 2026. Spherical Voronoi: Directional Appearance as a Differentiable Partition of the Sphere. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 22529–22538. Didan, K. 2015. MOD13C1 MODIS/Terra Vegetation Indices 16-Day L3 Global 0.05Deg CMG. NASA LP DAAC, https://doi.org/10.5067/MODIS/MOD13C1.006. Dollinger, J.; Robert, D.; Plekhanova, E.; Drees, L.; and Wegner, J. D. 2025. Climplicit: Climatic Implicit Embeddings for Global Ecological Tasks. In ICLR 2025 Workshop: Tackling Climate Change with Machine Learning. Hall, D. K.; and Riggs, G. A. 2015. MODIS/Terra Snow Cover 8-Day L3 Global 0.05Deg CMG, Version 6 (MOD10C2). NASA NSIDC DAAC, https://doi.org/10. 5067/MODIS/MOD10C2.006.
Karger, D. N.; Conrad, O.; Böhner, J.; Kawohl, T.; Kreft, H.; Soria-Auza, R. W.; Zimmermann, N. E.; Linder, H. P.; and Kessler, M. 2017. Climatologies at high resolution for the earth’s land surface areas. Scientific Data, 4: 170122. Karimzadeh, M.; Wang, Z.; and Crooks, J. L. 2025. Performance and generalizability impacts of incorporating location encoders into deep learning for dynamic PM2.5 estimation. GIScience & Remote Sensing, 62(1). Kattenborn, T.; Schiefer, F.; Frey, J.; Feilhauer, H.; Mahecha, M. D.; and Dormann, C. F. 2022. Spatially Autocorrelated Training and Validation Samples Inflate Performance Assessment of Convolutional Neural Networks. ISPRS Open Journal of Photogrammetry and Remote Sensing, 5: 100018. Klemmer, K.; Rolf, E.; Robinson, C.; Mackey, L.; and Rußwurm, M. 2025. Satclip: Global, general-purpose location embeddings with satellite imagery. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 4347–4355. Kratzert, F.; Nearing, G.; Addor, N.; Erickson, T.; Gauch, M.; Gilon, O.; Gudmundsson, L.; Hassidim, A.; Klotz, D.; Nevo, S.; Shalev, G.; and Matias, Y. 2023. Caravan — A global community dataset for large-sample hydrology. Scientific Data, 10: 61. Lang, N.; Jetz, W.; Schindler, K.; and Wegner, J. D. 2023. A high-resolution canopy height model of the Earth. Nature Ecology & Evolution, 7(11): 1778–1789. Loarie, S. R.; Duffy, P. B.; Hamilton, H.; Asner, G. P.; Field, C. B.; and Ackerly, D. D. 2009. The Velocity of Climate Change. Nature, 462(7276): 1052–1055. Lusk, D.; Wolf, S.; Svidzinska, D.; and Kattenborn, T. 2026. Global plant trait maps based on crowdsourced biodiversity monitoring and Earth observation - 1 km - All PFTs. Zenodo, https://doi.org/10.5281/zenodo.14646322. Mac Aodha, O.; Cole, E.; and Perona, P. 2019. Presence-Only Geographical Priors for Fine-Grained Image Classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Mai, G.; Janowicz, K.; Hu, Y.; Gao, S.; Yan, B.; Zhu, R.; Cai, L.; and Lao, N. 2022. A Review of Location Encoding for GeoAI: Methods and Applications. International Journal of Geographical Information Science, 36(4): 639–673. Mai, G.; Janowicz, K.; Yan, B.; Zhu, R.; Cai, L.; and Lao, N. 2020. Multi-Scale Representation Learning for Spatial Feature Distributions using Grid Cells. In The Eighth International Conference on Learning Representations (ICLR). Mai, G.; Lao, N.; He, Y.; Song, J.; and Ermon, S. 2023a. CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations. In Proceedings of the International Conference on Machine Learning (ICML). Mai, G.; Xuan, Y.; Zuo, W.; He, Y.; Song, J.; Ermon, S.; Janowicz, K.; and Lao, N. 2023b. Sphere2Vec: A GeneralPurpose Location Representation Learning over a Spherical Surface for Large-Scale Geospatial Predictions. ISPRS Journal of Photogrammetry and Remote Sensing, 202: 439–462. Mickisch, D.; Klemmer, K.; Teng, M.; and Rolnick, D. 2025. A Joint Space-Time Encoder for Geographic Time-Series
Data. In ICLR 2025 Workshop on Machine Learning Multiscale Processes. Milly, P. C. D.; Betancourt, J.; Falkenmark, M.; Hirsch, R. M.; Kundzewicz, Z. W.; Lettenmaier, D. P.; and Stouffer, R. J. 2008. Stationarity Is Dead: Whither Water Management? Science, 319(5863): 573–574. Muñoz-Sabater, J.; Dutra, E.; Agustí-Panareda, A.; Albergel, C.; Arduini, G.; Balsamo, G.; Boussetta, S.; Choulga, M.; Harrigan, S.; Hersbach, H.; Martens, B.; Miralles, D. G.; Piles, M.; Rodríguez-Fernández, N. J.; Zsoter, E.; Buontempo, C.; and Thépaut, J.-N. 2021. ERA5-Land: a state-ofthe-art global reanalysis dataset for land applications. Earth System Science Data, 13: 4349–4383. Ploton, P.; Mortier, F.; Réjou-Méchain, M.; Barbier, N.; Picard, N.; Rossi, V.; Dormann, C.; Cornu, G.; Viennois, G.; Bayol, N.; Lyapustin, A.; Gourlet-Fleury, S.; and Pélissier, R. 2020. Spatial Validation Reveals Poor Predictive Performance of Large-Scale Ecological Mapping Models. Nature Communications, 11(1): 4540. Rao, A.; Crasto, R.; Ooms, T.; Rolnick, D.; Klemmer, K.; and Rußwurm, M. 2026. Localized, High-resolution Geographic Representations with Slepian Functions. In Proceedings of the International Conference on Machine Learning (ICML). Rebain, D.; Jiang, W.; Yazdani, S.; Li, K.; Yi, K. M.; and Tagliasacchi, A. 2021. DeRF: Decomposed Radiance Fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14153–14161. Roberts, D. R.; Bahn, V.; Ciuti, S.; Boyce, M. S.; Elith, J.; Guillera-Arroita, G.; Hauenstein, S.; Lahoz-Monfort, J. J.; Schröder, B.; Thuiller, W.; Warton, D. I.; Wintle, B. A.; Hartig, F.; and Dormann, C. F. 2017. Cross-Validation Strategies for Data with Temporal, Spatial, Hierarchical, or Phylogenetic Structure. Ecography, 40(8): 913–929. Rolf, E.; Klemmer, K.; Robinson, C.; and Kerner, H. 2024. Position: Mission Critical – Satellite Data is a Distinct Modality in Machine Learning. In Proceedings of the 41st International Conference on Machine Learning (ICML). Rosemartin, A. H.; Denny, E. G.; Gerst, K. L.; Marsh, R. L.; Posthumus, E. E.; Crimmins, T. M.; and Weltzin, J. F. 2018. USA National Phenology Network Observational Data Documentation. Technical Report Open-File Report 2018-1060, U.S. Geological Survey. Rußwurm, M.; Klemmer, K.; Rolf, E.; Zbinden, R.; and Tuia, D. 2024. Geographic Location Encoding with Spherical Harmonics and Sinusoidal Representation Networks. In Proceedings of the International Conference on Learning Representations (ICLR). Sastry, S.; Khanal, S.; Dhakal, A.; and Jacobs, N. 2024. GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis. In IEEE/ISPRS Workshop: Large Scale Computer Vision for Remote Sensing (EARTHVISION). Shatwell, D. G.; Dave, I. R.; Swetha, S.; and Shah, M. 2025. GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Sitzmann, V.; Martel, J. N.; Bergman, A. W.; Lindell, D. B.; and Wetzstein, G. 2020. Implicit Neural Representations
with Periodic Activation Functions. In NeurIPS, volume 33, 7462–7473. Tancik, M.; Srinivasan, P. P.; Mildenhall, B.; Fridovich-Keil, S.; Raghavan, N.; Singhal, U.; Ramamoorthi, R.; Barron, J. T.; and Ng, R. 2020. Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains. In NeurIPS, volume 33, 7537–7547. Tseng, G.; Fuller, A.; Reil, M.; Herzog, H.; Beukema, P.; Bastani, F.; Green, J. R.; Shelhamer, E.; Kerner, H.; and Rolnick, D. 2025. Galileo: Learning Global and Local Features of Many Remote Sensing Modalities. In Proceedings of the 42nd International Conference on Machine Learning (ICML). USA National Phenology Network. 2012. USA National Phenology Network In-Situ Phenology Observation Data. Valavi, R.; Elith, J.; Lahoz-Monfort, J. J.; and GuilleraArroita, G. 2019. blockCV: An R Package for Generating Spatially or Environmentally Separated Folds for K-Fold Cross-Validation of Species Distribution Models. Methods in Ecology and Evolution, 10(2): 225–232. van der Werf, G. R.; Randerson, J. T.; Giglio, L.; van Leeuwen, T. T.; Chen, Y.; Rogers, B. M.; Mu, M.; van Marle, M. J. E.; Morton, D. C.; Collatz, G. J.; Yokelson, R. J.; and Kasibhatla, P. S. 2017. Global fire emissions estimates during 1997–2016. Earth System Science Data, 9: 697–720. Vivanco Cepeda, V.; Nayak, G. K.; and Shah, M. 2023. Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization. Advances in Neural Information Processing Systems, 36: 8690–8701. Wei, B.; Sastry, S.; Cher, D.; Xing, E.; and Jacobs, N. 2026. TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive. In European Conference on Computer Vision. Williams, J. W.; and Jackson, S. T. 2007. Novel Climates, NoAnalog Communities, and Ecological Surprises. Frontiers in Ecology and the Environment, 5(9): 475–482. Yin, Y.; Liu, Z.; Zhang, Y.; Wang, S.; Shah, R. R.; and Zimmermann, R. 2019. GPS2Vec: Towards Generating Worldwide GPS Embeddings. In Proceedings of the 27th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems, 416–419.