Drift Happens: An Empirical Study of Neural Architecture Robustness to Temporal Distribution Shift Robin Holzinger*
Riccardo Colletti*
Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, USA Extended version. Accepted at QCDS 2026; the proceedings version will appear in Springer LNCS.
arXiv:2607.05908v1 [cs.LG] 7 Jul 2026
Abstract
real-world data evolves over time, exhibiting temporal distribution shift that can degrade model performance in ways that standard held-out evaluation fails to capture. A model that achieves high accuracy on held-out data from the same time period may fail when deployed on future inputs.
Real-world data distributions evolve over time, inducing temporal distribution shift that can substantially degrade the reliability of deployed machine learning systems. However, the extent to which architectural choices and their associated inductive biases affect temporal robustness remains insufficiently understood.
While distribution shift is well-documented, less understood is how architectural choices influence a model’s robustness to temporal drift.
We present a systematic empirical comparison of temporal robustness across three heterogeneous, time-indexed domains encompassing image classification, multi-label text classification, and text regression tasks. Using a unified evaluation framework based on temporal drift matrices, we train models on cumulative historical data and evaluate their performance on both earlier and later time periods, thereby quantifying cross-temporal generalization. Our study spans model families ranging from simple multilayer perceptrons and convolutional networks to recurrent networks and pretrained Transformer-based encoders.
Do different inductive biases (the translation invariance of convolutions, the sequential modeling of recurrent nets, the attention of Transformers) lead to different rates of temporal degradation? Do frozen, pretrained encoders resist temporal drift better than models trained end to end? These questions have practical implications for model selection, yet systematic comparisons across architectures and domains remain scarce. This work investigates the temporal robustness of neural classifiers across three domains: image classification (Yearbook), text regression (Amazon Reviews), and multi-label text classification (arXiv). For each domain, we evaluate diverse architectures, from simple baselines to pretrained transformers, using a unified framework based on temporal drift matrices and provide qualitative explanations for model degradation via gradient saliency maps. Our contributions are threefold:
Collectively, the results show that architectural inductive biases systematically shape temporal robustness: models whose inductive biases lead them to exploit localized, highly discriminative features attain the highest in-distribution accuracy, yet those features are often the ones that change most over time, so these models degrade fastest, while pretrained encoders that draw on coarser, more stable representations drift more gradually. These observations offer practical guidance for selecting architectures for real-world systems subject to temporal drift.
• We provide a unified empirical assessment of temporal robustness across three long-range, time-indexed domains, enabling direct comparison of how neural architectures behave under temporal distribution shift. • We organize time-indexed evaluation in the spirit of Wild-Time [39] into temporal drift matrices, a compact representation that quantifies cross-temporal generalization by measuring performance when training on cumulative historical data and testing on both earlier and later time periods.
1. Introduction Machine learning models are typically trained under the assumption that training and test data are drawn from the same distribution. In practice, this assumption rarely holds:
• We systematically compare a broad spectrum of model families, from multilayer perceptrons, convolutional
* Both authors contributed equally.
1
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift Original
and residual networks, and recurrent networks to transformers trained on the data and frozen pretrained encoders, highlighting how architectural assumptions relate to degradation patterns across modalities and tasks.
Concept Drift
Covariate Drift
Label Drift
(Virtual Drift)
Taken together, our results characterize how performance deteriorates as the temporal gap between training and evaluation widens across domains and model classes. The features an inductive bias extracts within the training period are what lift in-distribution performance, yet they are the most tied to it and the first to degrade as the data drifts, so the very features that make a model accurate in distribution are the least robust over time. Frozen pretrained encoders, relying on coarser and more transferable representations, trade in-distribution accuracy for steadier degradation. These findings guide architecture selection in non-stationary real-world environments.
Figure 1: Drift types in a binary classification setting. Circles and stars indicate the label classes; the curve represents the decision boundary. Concept drift induces a change in the decision boundary, Covariate drift (virtual drift) changes the input distribution, and Label drift alters the relative class frequency [19]. In practice, these forms of drift rarely occur in isolation and often manifest in combination [2]. As a result, the temporal robustness of a model depends not only on the magnitude of drift but also on how its architectural assumptions and inductive biases interact with the evolving data distribution.
2. Background and Related Work
2.2. Model Robustness to Distribution Shift
2.1. Temporal Distribution Shift
Understanding how learning algorithms behave under distribution shift has become an important research direction. Early work approached robustness through domain adaptation [3] and out-of-distribution generalization [14], formalizing shift as a transition between a small number of discrete source and target domains. The introduction of large-scale benchmarks such as WILDS [23], which focuses on naturally occurring distribution shifts without explicit temporal indexing, and more recently Wild-Time [39], which explicitly models time-indexed data and temporal distribution shifts, has broadened this perspective by providing evaluation protocols that more closely reflect real-world deployment scenarios.
A central challenge in real-world machine learning systems is that the data-generating process is rarely stationary. As models are deployed over months, years, or decades, both inputs and label semantics may evolve, causing systematic discrepancies between distributions encountered during training and those observed at inference. Such temporal evolution, commonly referred to as temporal distribution shift, is pervasive across domains. Understanding the structure of this shift is essential for assessing and improving temporal robustness. However, prior work has largely focused on static or domain-level distribution shifts [3; 14; 23], with limited attention to drift that unfolds sequentially over time, particularly in non-generated, in-the-wild settings [39].
A complementary line of research has examined robustness through model diagnostics and monitoring. Methods for detecting drift onset by examining changes in latent representations or predictive uncertainty have shown promise in deep learning settings [29; 1]. Recent empirical analyses have characterized how concept drift manifests in practice, including its locality and temporal progression across largescale data streams [2]. At the systems level, work on datacentric and continual-learning infrastructures has explored how to maintain model quality over time through cost-aware retraining and pipeline orchestration [26; 27; 34; 4; 19].
Following established terminology in concept drift research [11], temporal distribution shift can be characterized along three complementary axes (Figure 1): Covariate shift occurs when the input distribution P (X) changes while the conditional P (Y | X) remains fixed. This type of shift commonly arises in natural data streams where observational setups, sociocultural conventions, or user behavior gradually evolve. Label shift refers to changes in the marginal label distribution P (Y ). Long-term textual or behavioral datasets frequently exhibit such imbalance drift as the prevalence of topics, categories, or rating patterns changes over time.
Despite these advances, robustness studies commonly focus either on a single architecture evaluated across multiple datasets or on a single dataset used to compare a narrow set of architectures. As a result, comparatively little is known about how architectural design choices interact with longrange temporal drift across heterogeneous modalities and tasks. This gap is particularly relevant given the diversity of inductive biases exhibited by modern neural models: convolutional networks encode locality and translation equivari-
Concept drift denotes changes in the conditional distribution P (Y | X), implying that identical inputs may correspond to different labels at different time points. This form of drift is particularly pronounced in domains where semantics or visual attributes evolve, such as historical portraits [12] or online platforms. 2
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
3.1.1. Y EARBOOK : A C ENTURY OF P ORTRAITS
ance, recurrent networks capture sequential structure [18; 5], and Transformer-based encoders rely on self-attention with minimal structural priors [35]. Likewise, large-scale pretraining in vision and language [8; 24] introduces representations shaped by broad historical data, yet their behavior under extended temporal drift remains poorly characterized.
The Yearbook dataset [13] contains 37,921 frontal portraits of American high-school seniors from 1905 to 2013, which we align and process as 3-channel 32 × 32 tensors. Although acquisition is largely standardized, stylistic attributes (e.g., hairstyles, clothing, accessories) vary substantially across decades. The dataset provides binary sex labels that are approximately balanced over time. It is a canonical benchmark for temporal shift: Wild-Time [39] uses a pre-/post-1970 split and reports marked out-of-distribution degradation, and Modyn [4] observes accuracy decay as the train–test time gap grows. We reproduce this trend (Fig. 2), consistent with covariate and concept drift.
The present study addresses this open question by comparing these architectural families under a shared temporal evaluation protocol and across multiple modalities, enabling a controlled analysis of how model design influences robustness under real-world temporal distribution shift. 2.3. Robustness Evaluation vs. Temporal Adaptation This work evaluates the inherent temporal robustness of neural architectures under distribution shift, measuring how different model families, trained only on cumulative historical data, perform as the temporal gap between training and evaluation widens. This isolates the robustness arising from architectural design and pretraining alone, prior to any adaptive intervention.
3.1.2. A MAZON R EVIEWS 2023: E- COMMERCE The Amazon Reviews dataset [20] comprises 571.54 million reviews across 33 product categories, spanning May 1996 to September 2023. Each review has a timestamp, star rating, and free-text content. We cast a review-level sentiment regression task, predicting the 1–5 rating from the text, which exhibits strong covariate and concept drift due to evolving language, consumer behavior, and platform usage. We focus on seven categories, restrict to 2014–2023, and draw a stratified sample of 300,000 reviews.
A complementary line of research instead investigates how models can adapt once drift is detected, through scheduled or cost-aware retraining [26; 27], continual-learning pipelines [34; 4; 19], or accuracy-aware data maintenance [38], addressing when to retrain, how much data to incorporate, and how to trade performance against cost. Our study is orthogonal to that literature and provides a foundation on which such adaptation strategies can be designed, evaluated, and compared.
3.1.3. AR X IV: S CIENTIFIC D ISCOURSE The arXiv dataset [7] defines a multi-label task over 2,866,787 title–abstract records annotated with 176 subject categories. We concatenate titles and abstracts, keep seven leaf categories (Section E), and use 2000–2025 submissions whose categories fall within them (papers may carry several). Category mix and terminology shift, inducing label and covariate drift [19]. Wild-Time [39] reports roughly 20% degradation under temporal splits for a related arXiv task, and the gap is not substantially closed by domain generalization or continual learning methods, motivating our architectural comparison.
3. Methods and Experimental Approach Temporal distribution shift manifests differently across modalities, tasks, and time scales, yet existing empirical studies typically vary a single axis at a time, leaving open how architectural design, label structure, and drift mechanism jointly shape temporal robustness (Section 2). Our setup targets exactly this cross-cutting comparison: we combine three long-range, time-indexed datasets with a diverse suite of neural architectures spanning multiple inductive biases, all evaluated under the unified temporal protocol of Section 4 (temporal drift matrices), so that architectural differences, rather than differences in splitting or evaluation, drive the observed robustness patterns.
3.2. Model Implementation We evaluate architectures spanning different inductive biases to understand how model design affects temporal robustness. Rather than covering the full architecture landscape, we select families that span the spectrum of inductivebias strength: from simple baselines without structural assumptions that establish lower bounds on performance, to modern architectures that incorporate structural priors, to pre-trained models that leverage large-scale external data. This spread is what lets us attribute robustness differences to the priors themselves.
3.1. Datasets Our empirical analysis spans three qualitatively distinct temporal scenarios, chosen to cover different modalities, tasks, and sources of distribution shift. Each dataset captures multiple decades of real-world temporal evolution, making it suitable for cross-temporal evaluation.
3
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
3.2.1. I MAGE C LASSIFICATION
number and width of its filters. The recurrent models read the sequence token by token, as a single-layer bidirectional GRU [5], a single-layer bidirectional LSTM [18], and a twolayer bidirectional LSTM with attention. The Transformer encoders replace recurrence with self-attention [35] and learnable positional embeddings, ranging from 1 to 5 layers and 4 to 6 heads.
For the Yearbook dataset, we evaluate four model families trained on the data, each at three sizes (small, medium, large), for 12 architectures, complemented by 9 pretrained vision encoders used as frozen backbones. The four families differ in the spatial prior they encode. At one extreme, the multilayer perceptron (MLP) flattens the image and applies only fully connected layers, imposing no spatial structure. The convolutional network (CNN) builds in locality and translation equivariance through convolutional filters with batch normalization and pooling, and the residual network (ResNet) extends it with skip connections [15] that ease optimization at greater depth. At the other extreme, the vision Transformer (ViT) drops convolution for self-attention over patches [9], a far weaker spatial prior, with a small patch size suited to the 32 × 32 inputs. Across all four, the small, medium, and large variants scale depth and width.
Finally, we assess transfer learning from pretrained language encoders. We consider frozen encoders with a light head trained on the pooled output: BERT [8], RoBERTa [24], DeBERTa-v3 [17], ELECTRA [6], MPNet [33], ModernBERT [37], and DistilBERT [31] on both text tasks, plus MiniLM-L6 [36] on arXiv. As with the vision encoders, freezing isolates pretrained representations from fine-tuning.
4. Evaluation Framework Standard held-out evaluation measures performance on data from the training distribution, which under temporal shift can look strong even as the model fails on future data. We therefore measure cross-temporal generalization explicitly.
Finally, we assess transfer learning using pretrained vision encoders trained on large-scale curated or web-scale datasets. The underlying hypothesis is that representations learned from diverse data may exhibit stronger temporal robustness than features learned solely from historical portraits. We consider 9 pretrained models. The selfsupervised DINOv2-S [28] and DINOv3-S [32] are pretrained on large curated image collections; CLIP-B32 [30] and SigLIP-B [40] use contrastive image-text pretraining, the latter with a sigmoid loss; ConvNeXt-S [25], ResNet50-IN [15], and ViT-S16-IN21k [9] are trained with supervised ImageNet labels, the last on ImageNet-21k; and MAE-B [16] and EVA02-B [10] are pretrained with masked image modeling. In all cases, we freeze the pretrained backbone and train only a linear classification head, isolating the contribution of pretrained representations from the effects of fine-tuning dynamics.
4.1. Temporal Drift Matrices Let D = {(xi , yi , ti )}N i=1 denote a dataset where each example is associated with a timestamp ti . We divide the timeline into K disjoint intervals T = {T1 , . . . , TK }, where end Tk = [tstart k , tk ). Let Dk = {(x, y, t) ∈ D : t ∈ Tk } denote the subset of data from interval k. For training, we construct cumulative datasets D≤k = Sk j=1 Dj containing all data up to and including interval k. A model fk trained on D≤k has access to historical data over time tend k but no knowledge of future periods. This cumulative strategy reflects realistic deployment scenarios in which models are periodically retrained on all historical data.
3.2.2. T EXT M ODELS
The temporal drift matrix M ∈ RK×K captures crosstemporal generalization:
For the Amazon Reviews and arXiv datasets, we evaluate model families that differ in their inductive biases for representing textual structure. The same architectures serve both datasets, differing only in output layer and loss, with Amazon Reviews using regression under a weighted mean squared error and arXiv multi-label classification under a weighted binary cross-entropy.
Mij = perf(fi , Dj )
(1)
where perf(·, ·) denotes a performance metric (accuracy, macro AUC, or balanced MSE depending on the task). Entry Mij measures how well a model trained on data through period i performs on data from period j.
All four families operate on cached RoBERTa token embeddings. The feed-forward baseline (FFN) averages them into a single 768-dimensional vector and passes it through a small MLP head that discards order entirely, with variants scaling the hidden width from 128 to 2048. The convolutional model (TextCNN) [21] applies one-dimensional convolutions to capture local n-gram patterns, scaling the
The structure of M reveals aspects of temporal robustness. Following our plotting convention, rows are the training cutoff (“trained up to”) and columns the evaluation period (“evaluated on”), with time increasing upward and to the right from a lower-left origin. The diagonal Mii is in-distribution performance on data held out from the training period; the upper-left (j < i) is held-out performance 4
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
on earlier in-training periods; and the lower-right (j > i) is forward generalization to data unseen during training. This lets us read how performance degrades (or improves) as the temporal gap between training and evaluation widens.
belong to multiple subject categories. We keep only seven leaf categories as the label space, retaining papers that carry at least one of them. The label distribution remains skewed, with negative examples substantially outnumbering positives for each category. We therefore use a weighted binary cross-entropy with logits, applied independently to each of the C = 7 categories:
4.2. Temporal Splitting Strategy Each slice Dk is partitioned once into a stratified training split Dktrain (70%) and a held-out test split Dktest (30%), shared across all models via a fixed split seed. For in-distribution evaluation (when j ≤ i), we use only the held-out test split of Dj to prevent data leakage:
( (i,j)
Deval =
Djtest Djtrain ∪ Djtest
n
LBCE = −
C
1 XX wc yic log(σ(zic )) n i=1 c=1 + (1 − yic ) log(1 − σ(zic )) ,
(2)
where zic denotes the logit for class c, σ(·) is the sigmoid function, and yic ∈ {0, 1} is the binary label. The class weights wc are defined as
if j ≤ i (in-distribution) if j > i (out-of-distrib.)
|{i : yic = 0}| , |{i : yic = 1}|
wc = For out-of-distribution evaluation on future time slices, we use all available samples from that period to maximize statistical power, since by definition none of this data was seen during training. Evaluating on periods that overlap the training data (j ≤ i) is deliberate: these held-out entries verify that a model retains performance across the historical periods it was trained on and provide the in-distribution reference from which forward decay is measured.
(3)
i.e., the ratio of negative to positive examples for each category. This reweighting compensates for label imbalance by amplifying the contribution of rare positive labels, preventing the model from achieving deceptively high accuracy by predicting the all-zero vector. For regression on Amazon Reviews, we train models to predict the star rating using a weighted mean squared error:
4.3. Training Protocol
n
LWMSE =
The image models are trained with Adam [22] and the text models with its decoupled-weight-decay variant AdamW, a standard choice that behaves robustly across heterogeneous architectures; fixing one optimizer per modality avoids permodel optimizer tuning as a confound. Learning rates are task-specific, with exact hyperparameters pinned in versioned experiment presets. Training proceeds for a fixed number of epochs, and model selection is based on the final checkpoint. Every configuration is trained under multiple random seeds, five on Yearbook and three on the text tasks (Section A), and all results average over seeds.
1X wy (yi − ŷi )2 , n i=1 i
(4)
where yi ∈ {1, 2, 3, 4, 5} is the true rating and ŷi is the prediction. The rating-specific weights wr are defined as wr =
n , nr
(5)
with n the total number of samples and nr the number of samples with rating r. Because the empirical rating distribution is heavily skewed toward high scores, an unweighted MSE would be dominated by the majority class and largely ignore rare low-rating events. The inverse-frequency weighting counteracts this imbalance, ensuring that errors on minority ratings remain visible in the objective and that temporal degradation in performance cannot be explained solely by changes in the prevalence of positive reviews.
We adopt a cumulative temporal training strategy: for each slice Tk we train a model fk on D≤k , all examples observed up to Tk , yielding a sequence {f1 , . . . , fK } that each represent the best model obtainable from the data available at that point in time. All checkpoints are stored and evaluated on every slice to construct the full temporal drift matrix M .
4.4. Evaluation Metrics
The loss follows the task structure and is tailored to address label imbalance.
To populate each entry of the drift matrix M , we require a scalar performance metric per train-test time pair. We match each metric to the task’s label structure: accuracy for Yearbook, whose binary labels are approximately balanced (Section B.1); balanced MSE (MSEbal ) for Amazon Reviews, which reweights its skewed rating distribution (Section B.2); and macro AUC (AUC) for arXiv, whose
For binary classification on Yearbook, we minimize standard two-class cross-entropy on logits and labels, the maximum-likelihood objective for a two-class softmax model. For multi-label classification on arXiv, each paper may 5
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
5. Results
imbalanced multi-label categories require a threshold-free, per-class average (Section B.3). The two classification tasks thus receive different metrics because their label structures differ; the drift protocol itself is identical.
Each model is summarised through its drift matrix by the in-distribution, forward, and decay scores of Section 4.5, namely its performance on the period it was trained through, its mean performance on the later periods held out from training, and the difference between the two. How far a model degrades between them is governed by two factors that recur across the three domains, the strength of its inductive bias and whether it relies on a frozen pretrained encoder.
4.5. Summary Statistics of the Drift Matrix In the drift matrix M , the row index i is the training cutoff and the column index j is the evaluation period, so the entry Mij is the performance of a model trained on all data up to period i when it is tested on period j. Three numbers summarize the matrix, each averaged only over the cells that are filled, since a train-test pair with no completed run leaves its cell empty. The in-distribution score is the average of the diagonal, ID(M ) = meani Mii ,
5.1. Image Classification On Yearbook, the model families separate sharply in distribution, along the diagonal of the cohort-mean matrix in the Yearbook panel of Figure 2. MLP-S, which flattens the image into a vector and imposes no spatial structure, sits at about 69.3%, while the CNNs and ResNets reach near 92.8% for CNN-M and CNN-L (Table 5). What sets these apart is their inductive bias toward locality and translation equivariance, by which they build features from small local regions of the image and recognise a pattern wherever it appears, and this lets them exploit the cues that most sharply separate male from female portraits within a given period. The ViTs and frozen encoders fall between these two ends.
(6)
a model’s performance on the same period it was trained through. The future score averages the cells with j > i, where the evaluation period falls after the training cutoff, Fut(M ) = meanj>i Mij .
(7)
The decay is the difference between the two, oriented so a larger value always means worse temporal robustness, since accuracy and AUC are better when high while MSEbal is better when low, (
Temporal robustness reverses this ordering, and we read it from the decay, the accuracy a model loses once the evaluation year moves past its training cutoff (Table 5). The CNNs and ResNets, strongest in distribution, decay the most, by 13.7 and 13.9 points for CNN-L and CNN-M, so their accuracy on future years falls to about 79.0%. MLP-S is the steadiest model of all, with only 7.1 points of decay, though from its lower starting accuracy of 62.3%. The frozen pretrained encoders form a third group, more stable than the CNNs but less accurate to begin with, decaying by 7.9 to 10.4 points, with the self-supervised DINOv3-S steadiest among them and the supervised ImageNet-21k backbone least.
ID(M ) − Fut(M ) Fut(M ) − ID(M )
(accuracy, AUC), (MSEbal ). (8) so a positive decay is a loss of performance on future periods and a negative one, which is uncommon, a gain. Together, future score and decay order every model from most to least temporally robust. Dec(M ) =
Beyond these whole-matrix scores, we look at how robustness depends on the training time itself. Fixing one cutoff i, a single row of the matrix, we pair its diagonal cell Mii with the average over the later periods in that row (meanj>i Mij ) and the decay between them, reading these at a few cutoffs spread evenly across the timeline (the last one is left out, as it has no future). Averaging the same quantities within each architecture family lets us compare the families directly.
What lifts a model in distribution is also what undermines it over time. The features a strong inductive bias extracts to separate the classes raise its in-distribution accuracy, but they are the most specific to the years it was trained on, and the first to lose their meaning once hairstyles, clothing, and image quality drift. The sharper a model’s in-distribution lead, the faster it erodes as the data ages. Robustness on Yearbook is therefore not a property an architecture optimises on its own, but the other side of fitting a single period too closely, closer to overfitting across time than across samples (Section C).
To place each model against the others, we average the matrices over the cohort C of models into the cohort-mean matrix 1 X (m) M̄ij = Mij , (9) |C| m∈C
defined at the cells every model fills; each model’s deviation (m) (m) is ∆ij = Mij − M̄ij . 6
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Evaluation year
Macro AUC (%)
20 17
Training year
90
20 08
20 17
0.8
95
2 20 000 00 20 08 20 17 20 25
20 20
1.0
Balanced MSE
20 23
20 25
arXiv
2 20 014 14 20 17 20 20 20 23
Accuracy (%)
19 79 19 44
60
1 19 90 05 5 19 44 19 79 20 13
Training year
80
Training half-year
Amazon Reviews
20 13
Yearbook
Evaluation half-year
Evaluation year
Figure 2: Cohort-mean drift matrices, one panel per domain. Each cell (i, j) is the mean performance across all models of that dataset when trained on data through period i (row) and evaluated on period j (column): accuracy for Yearbook, balanced MSE for Amazon Reviews (lower is better), and macro AUC for arXiv. The dashed diagonal marks indistribution evaluation; the lower-right region of each panel is forward generalization to unseen future periods. Vertical white bands mark periods with insufficient samples.
CNN-M CNN-S ConvNeXt-S DINOv2-S DINOv3-S EVA02-B MAE-B MLP-L MLP-M MLP-S ResNet50-IN ResNet-L ResNet-M ResNet-S SigLIP-B ViT-L ViT-M ViT-S ViT-S16-IN21k
80 Accuracy (%)
5.1.1. S ALIENCY M APS
CLIP-B32 CNN-L
90
70
60
50 0
20
40 60 Years since training
80
The gradient saliency maps in Figure 3 show where this fragility comes from. Extending the training window from 1950 to 1970 sharpens where the convolutional models look: the CNN tightens from a diffuse scatter over the cheeks and mouth to a compact, near-symmetric pair of bright spots at the eyes, the detail that most sharply tells the classes apart, and the ResNet shifts the same way while keeping a broader spread across the central face. The MLP, without that bias, spreads its attribution across the whole frame at either cutoff, the background included. The same precision that lets the convolutional models read this discriminative detail binds them to it: those localized features are the most specific to the training period and the most exposed to drift, while the MLP’s blunter reading is less tied to any era.
100
Figure 4: Forgetting curves on Yearbook. Each curve is one model’s mean accuracy as the gap between its training year and the evaluation year grows, averaged over training years and seeds. A zero gap is in-distribution, and the slope is the rate of forgetting (Section C.7).
5.2. Text Models Each text model is a light head trained on cached, frozen RoBERTa embeddings (Section 3), a shared representation over which the families differ only in inductive bias. Amazon Reviews is a rating regression scored by balanced MSE, where lower is better, and arXiv is a multilabel classification scored by macro AUC.
The forgetting curves in Figure 4 show this playing out year by year. Each curve plots a model’s accuracy against the gap between its training cutoff and the evaluation year, so reading from left to right traces how quickly it forgets (Section C.7). The strongly biased networks start far above the rest, near 90% at a zero gap, and their curves descend the steepest. As the gap widens the curves draw together and then cross, so the families that led in distribution lose their edge on the most distant years, and every model settles near two-class chance.
5.2.1. A MAZON R EVIEWS On Amazon Reviews the recurrent and Transformer models fit the training reviews most closely, since both read the wording in context, the recurrent networks token 7
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift Train ≤1950 Eval 1960
Eval 1980
Train ≤1970 Eval 2000
Eval 1960
Eval 1980
Eval 2000
mlp_l
Saliency
resnet_s
cnn_l
high
low
Figure 3: Gradient saliency maps on Yearbook for CNN-L, ResNet-S, and MLP-L, each trained through 1950 (left) and 1970 (right) and evaluated on later years. Each panel averages gradient saliency over the selected portraits from the indicated evaluation year rather than a single example. ranking inverts, and the models that fit the reviews most tightly end among the least accurate, while the order-free FFN, never sharp, stays the steadiest.
by token and the Transformers through self-attention, and so capture how a review’s words compose into its rating. That fit shows up as the lowest in-distribution error, down to 0.680 balanced MSE for BiGRU-S, with the Transformers alongside them at 0.686 for TX-S (Figure 2, Table 13). The FFN, which averages the embeddings and discards word order, sits higher at about 0.761, and the frozen encoders higher still, between 0.806 and 1.022.
5.2.2. AR X IV On arXiv a stronger inductive bias yields no in-distribution advantage, and so costs no temporal robustness. Sorting a paper into its subject categories is close to recognising its topic, and the topic is already encoded in the frozen RoBERTa embeddings every model is built on, so no architecture finds structure the others miss. The feed-forward, convolutional, recurrent, and Transformer families therefore reach almost the same in-distribution macro AUC, between about 97.2% and 98.1%, within a point of one another (Figure 2, Table 21).
Temporal robustness runs the other way, and the models that fit the reviews most tightly are the ones whose error grows fastest. The recurrent networks decay the most, by up to 0.128 balanced MSE, so their error on future reviews climbs to about 0.824, while the FFN is the steadiest of the trained models, gaining only 0.075 to 0.089. The decay is sharpest for models trained on the earliest reviews, where the recurrent family worsens by 0.242, and it narrows toward the recent cutoffs as fewer years remain ahead (Section D). The frozen encoders again sit at higher error but drift less than the trained models, and DeBERTa-v3 is the most stable model anywhere in the study, at 0.043 of decay.
Robustness is just as uniform, and each family loses a similar small 2.7 to 3.4 points going forward, with the frozen encoders in the same band apart from DeBERTa-v3 and ELECTRA, which fall further at 6.5 and 7.3 points (Section E). Where the bias gains nothing in distribution it forms no period-specific features to lose, so no family decays faster than the rest. The later years stay close to the training distribution: the backbone learned this vocabulary before any model saw it, leaving temporal shift little to take away.
What makes these models fit the reviews so well is also what later undermines them. The way they compose the wording into a rating is the most specific to how reviews were written in the training years, and the first to lose its meaning as the language shifts. Far enough forward the in-distribution 8
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
6. Conclusion and Future Work
such as Modyn [4].
Across image classification, text regression, and multi-label text classification, in-distribution accuracy turns out to neither guarantee nor predict temporal robustness. What governs the rate of degradation is instead the strength of a model’s inductive bias and whether it relies on a frozen pretrained encoder. A stronger bias extracts the most discriminative, period-specific features and leads in distribution, yet those same features are the first to lose their meaning as the data drifts, so the architectures that fit the training period most tightly are the ones that degrade fastest. Frozen pretrained encoders occupy a different regime, conceding in-distribution accuracy in return for steadier behaviour over time, and on arXiv, where the task is already solved by the pretrained representation and the bias buys no advantage, no family decays faster than the rest. The practical reading is that the model with the best held-out score at training time is often the least robust once deployed, so architecture selection should weigh the expected horizon between retraining cycles, not in-distribution accuracy alone.
6.2. Future Work These limitations point to several extensions. The comparison can be widened to a broader range of inductive biases, to capacity-matched pairs that disentangle bias from scale, and to additional datasets, modalities, and drift mechanisms, and carried to longer temporal spans with the full corpora and richer label spaces, where language drift should be more pronounced and the gap between architectures wider. The drift matrices can also be paired with scheduled or drift-triggered retraining under explicit compute budgets, measuring not only how much a model decays but how cheaply that decay can be undone [27]. Most of all, moving from description to mechanism, by tying inductive bias and pretraining back to the specific features that drift, would turn the observed regularity into an account of why temporal robustness behaves as it does. 6.3. Availability Code, experiment definitions, and paper-generation scripts are available at github.com/learning-mechanisms/drifthappens; public run histories and matrix artifacts at wandb.ai/drift-happens/drift-happens. The extended preprint and drift-happens.org retain the complete per-model drift-matrix galleries. Code is Apache-2.0; paper, figures, and documentation are CC BY 4.0 where we hold the rights, and external datasets and models keep their original licenses.
6.1. Limitations Our comparison covers the most common inductive biases, convolutional, recurrent, and attention-based, against simple baselines and frozen encoders, but many architectures remain untested and the study is best read as a first systematic pass rather than an exhaustive one. It spans three datasets across two modalities, and broadening it to further domains and other forms of distribution shift would show how widely the pattern holds. The text tasks are the most constrained, using a narrow label space, a five-point rating for Amazon Reviews and seven categories for arXiv, spanning only about a decade and a quarter of a century against a full century for Yearbook, and running on a stratified subsample rather than the complete corpora. Because too few examples fall within each time slice to train a text encoder from scratch, the text models also share a single frozen RoBERTa representation, so their absolute scores should be read in that light. We average over a small fixed seed set rather than conducting formal significance tests, and hyperparameter choices and seed variance may still affect fine-grained rankings; we therefore emphasize family-level patterns over individual placements. The families are also not capacitymatched. The small, medium, and large variants within each family expose scale effects, and on Yearbook the smallest MLP, CNN, and ResNet variants decay less than their larger siblings, so inductive bias and capacity remain partially confounded. The analysis is also descriptive rather than mechanistic. We measure how robustness varies with inductive bias and pretraining without isolating the precise cues responsible, and we evaluate inherent robustness under cumulative training rather than any adaptive policy, which leaves the question of retraining orchestration to systems
References [1] Samuel Ackerman, Orna Raz, Marcel Zalmanovici, and Aviad Zlotnick. Automatically detecting data drift in machine learning classifiers, 2021. URL https: //arxiv.org/abs/2111.05672. [2] Gabriel J. Aguiar and Alberto Cano. A comprehensive analysis of concept drift locality in data streams, 2023. URL https://arxiv.org/abs/2311. 06396. [3] Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan. A theory of learning from different domains. Machine Learning, 79(1):151–175, 2010. [4] Maximilian Böther, Ties Robroek, Viktor Gsteiger, Robin Holzinger, Xianzhe Ma, Pınar Tözün, and Ana Klimovic. Modyn: Data-centric machine learning pipeline orchestration. Proc. ACM Manag. Data, 3 (1), February 2025. doi: 10.1145/3709705. URL https://doi.org/10.1145/3709705. [5] Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger 9
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734. Association for Computational Linguistics, 2014. doi: 10.3115/v1/D14-1179.
[14] Ishaan Gulrajani and David Lopez-Paz. In search of lost domain generalization. In International Conference on Learning Representations, 2021. [15] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
[6] Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. ELECTRA: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations (ICLR), 2020.
[16] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
[7] Cornell University. arxiv dataset, 2024. URL https: //www.kaggle.com/dsv/7548853.
[17] Pengcheng He, Jianfeng Gao, and Weizhu Chen. DeBERTaV3: Improving DeBERTa using ELECTRAstyle pre-training with gradient-disentangled embedding sharing. In International Conference on Learning Representations (ICLR), 2023.
[8] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACLHLT), pages 4171–4186. Association for Computational Linguistics, 2019. doi: 10.18653/v1/N19-1423.
[18] Sepp Hochreiter and Jürgen Schmidhuber. Long shortterm memory. Neural Computation, 9(8):1735–1780, 1997. doi: 10.1162/neco.1997.9.8.1735. [19] Robin Holzinger. An analysis of drift- and cost-aware ml retraining triggering policies in modyn. Bachelor’s thesis, Technical University of Munich, September 2024. Supervisors: Prof. Dr. Viktor Leis., Jana Vatter, Prof. Dr. Ana Klimović, Maximilian Böther.
[9] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
[20] Yupeng Hou, Jiacheng Li, Zhankui He, An Yan, Xiusi Chen, and Julian McAuley. Bridging language and items for retrieval and recommendation. arXiv preprint arXiv:2403.03952, 2024. [21] Yoon Kim. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1746–1751. Association for Computational Linguistics, 2014. doi: 10.3115/v1/D14-1181.
[10] Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. EVA-02: A visual representation for neon genesis. Image and Vision Computing, 149:105171, 2024. [11] João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia. A survey on concept drift adaptation. ACM Computing Surveys, 46(4):1–37, 2014. doi: 10.1145/2523813.
[22] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
[12] Shiry Ginosar, Kate Rakelly, Sarah Sachs, Brian Yin, and Alexei A Efros. A century of portraits: A visual historical record of american high school yearbooks. In IEEE International Conference on Computer Vision Workshops, pages 1–7, 2015.
[23] Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. WILDS: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning, pages 5637–5664. PMLR, 2021.
[13] Shiry Ginosar, Kate Rakelly, Sarah Sachs, Brian Yin, Crystal Lee, Philipp Krahenbuhl, and Alexei A. Efros. A century of portraits: A visual historical record of american high school yearbooks, 2019.
[24] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke 10
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Mairal, Hervé Jégou, Patrick Labatut, and Piotr Bojanowski. DINOv3. arXiv preprint arXiv:2508.10104, 2025.
[25] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A ConvNet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
[33] Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and TieYan Liu. MPNet: Masked and permuted pre-training for language understanding. In Advances in Neural Information Processing Systems (NeurIPS), 2020. [34] Huangshi Tian, Minchen Yu, and Wei Wang. Continuum: A platform for cost-aware, low-latency continual learning. In Proceedings of the ACM Symposium on Cloud Computing, SoCC ’18, page 26–40, New York, NY, USA, 2018. Association for Computing Machinery. ISBN 9781450360111. doi: 10.1145/3267809.3267817. URL https://doi. org/10.1145/3267809.3267817.
[26] Ananth Mahadevan and Michael Mathioudakis. Costeffective retraining of machine learning models, 2023. [27] Ananth Mahadevan and Michael Mathioudakis. Cost-aware retraining for machine learning. Knowledge-Based Systems, 293:111610, 2024. ISSN 0950-7051. doi: https://doi.org/10. 1016/j.knosys.2024.111610. URL https: //www.sciencedirect.com/science/ article/pii/S0950705124002454.
[35] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.
[28] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, ShangWen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jégou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research (TMLR), 2024.
[36] Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. MiniLM: Deep selfattention distillation for task-agnostic compression of pre-trained transformers. In Advances in Neural Information Processing Systems (NeurIPS), 2020. [37] Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, and Iacopo Poli. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. arXiv preprint arXiv:2412.13663, 2024.
[29] Stephan Rabanser, Stephan Günnemann, and Zachary Lipton. Failing loudly: An empirical study of methods for detecting dataset shift. In Advances in Neural Information Processing Systems, volume 32, 2019. [30] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), 2021.
[38] Sarah Wooders, Xiangxi Mo, Amit Narang, Kevin Lin, Ion Stoica, Joseph M. Hellerstein, Natacha Crooks, and Joseph E. Gonzalez. Ralf: Accuracy-aware scheduling for feature store maintenance. Proc. VLDB Endow., 17(3):563–576, nov 2023. ISSN 2150-8097. doi: 10.14778/3632093.3632116. URL https:// doi.org/10.14778/3632093.3632116.
[31] Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
[39] Huaxiu Yao, Caroline Choi, Bochuan Cao, Yoonho Lee, Pang Wei Koh, and Chelsea Finn. Wild-time: A benchmark of in-the-wild distribution shift over time. In Advances in Neural Information Processing Systems, 2022.
[32] Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea
[40] Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 11
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Appendix
B.1. Yearbook
A. Reproducibility
We slice the 1905–2013 timeline into one-year intervals, keeping the K = 104 years with enough samples. The task is binary classification with roughly balanced classes, so we score it with standard accuracy, the fraction of portraits classified correctly. Balance is what makes accuracy trustworthy here: when neither class dominates, a model cannot earn a good score by always predicting the same label, so accuracy rises and falls only as the model genuinely classifies more or fewer cases correctly. Higher values are better.
Experiments are defined as versioned preset snapshots and run through the repository command-line interface in a Pixi-pinned environment; run metadata records the git state and pixi.lock hash, and regenerated paper and website assets are checked against a committed checksum manifest. Each experiment is a staged run for one dataset, architecture, and seed: the training stage fits all cumulative temporal checkpoints, and the evaluation stage scores them on every time slice to form the drift matrix. We use seeds 0–4 for Yearbook and 0–2 for Amazon Reviews and arXiv, averaging over them. Public run histories and per-run matrix artifacts are available in the Weights & Biases project at https://wandb.ai/ drift-happens/drift-happens.
B.2. Amazon Reviews We slice the 2014–2023 timeline into half-year intervals, giving K = 20 slices. The task is regression: the model predicts a star rating from one to five and is scored by squared error. The difficulty is that the ratings are strongly skewed toward five-star reviews, so a single mean squared error taken over all reviews would mainly measure performance on the majority and would barely react to the rarer low ratings. To keep every rating level visible, we first compute the mean squared error separately within each rating and then average those per-rating values equally over the rating levels present,
Table 1 summarizes the measured compute on NVIDIA A100 GPUs, summing the W&B _runtime of each finished train and evaluation stage. GPU-hours are successfulstage wall times under one GPU per run; the four-GPU node-hour column divides these by four for ideal packed execution, excluding queueing, data preprocessing, failed attempts, and scheduling idle time. Table 1: Approximate measured compute for the conference experiment campaign. Dataset
MSEbal =
1 X MSEr , |R| r∈R
MSEr =
1 X (yi −ŷi )2 , |Sr | i∈Sr
Model-seed configs Train GPU-h Eval GPU-h 4-GPU node-h
Yearbook Amazon Reviews arXiv
105 57 60
23.7 116.6 272.3
9.3 136.7 101.9
8.2 63.3 93.5
Total
222
412.6
247.9
165.1
where Sr = {i : yi = r} is the set of reviews whose true rating is r and R = {r : |Sr | > 0} is the set of rating levels present in the slice (at most the five levels one to five). Because every present level contributes equally, a model that began to drift on the scarce one- and two-star reviews would reveal it here. Unlike the other two metrics, lower values are better.
B. Per-Dataset Protocol and Metrics The drift-matrix protocol of Section 4 is shared across the three datasets, but each instantiates it at its own time granularity and is scored with the metric matched to its label structure (Table 2). For each task we require a scalar performance measure that is both aligned with the task and robust to the class imbalance present in that dataset; a metric dominated by the label distribution would conflate shifts in the data with changes in the model.
B.3. arXiv We slice the 2000–2025 timeline into one-year intervals, giving K = 26 slices. The task is multi-label: a paper can belong to several of the C = 7 subject areas at once, and those areas are very unevenly populated. This raises two problems. First, fixing a single decision threshold is arbitrary and tends to favour the frequent subjects, so a threshold-based score such as accuracy is misleading. Second, an average that weights papers equally is again dominated by the common subjects. We address the first by scoring each subject with the Area Under the ROC Curve, which summarises performance over all thresholds at once: AUCc is the probability that a randomly chosen paper carrying subject c is given a higher score for that subject than a randomly chosen paper that does not carry it. We address the second by averaging the per-subject values uniformly, so each subject counts the
Table 2: Per-dataset instantiation of the drift-matrix protocol: time span, slice granularity, number of slices K, and primary metric. Dataset
Span
Slice
K
Metric
Yearbook Amazon Reviews arXiv
1905–2013 2014–2023 2000–2025
yearly half-yearly yearly
104 20 26
accuracy balanced MSE macro AUC
12
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
same regardless of how many papers it contains, C
1 X AUC = AUCc , C c=1
Z 1 AUCc =
TPRc (t) dFPRc (t), 0
where TPRc (t) and FPRc (t) denote the true- and falsepositive rates for subject c at threshold t. Higher values are better, and a model that ranked every paper correctly within every subject would reach 1.
13
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
C. Yearbook – Drift Matrices The cohort-mean and per-model deviation matrices shown here, and the in-distribution, future, and decay quantities tabulated below, are defined in Section 4.5.
2013 2009 2000 90 1991
80
1973 1964 1955
70
Accuracy (%)
Training year
1982
1946 1937
60
1928 1915 50
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
19 05
1905
Evaluation year
Figure 5: Cohort-mean Accuracy matrix M̄ over the Yearbook models. Cell (i, j) is the mean across those models of the score from training through slice i and evaluating on slice j.
C.1. Model Roster Table 3: Yearbook: models trained from scratch. Model
Family
Parameters
ViT-S CNN-S ResNet-S MLP-S ResNet-M MLP-M CNN-M ViT-M MLP-L ResNet-L CNN-L ViT-L
ViT CNN ResNet MLP ResNet MLP CNN ViT MLP ResNet CNN ViT
75k 94k 98k 98k 400k 410k 542k 545k 2.1M 2.1M 2.1M 2.2M
Table 4: Yearbook: frozen pretrained encoders, with trainable head and total parameters.
14
Model
Family
Trainable
Total
DINOv3-S ViT-S16-IN21k DINOv2-S ResNet50-IN ConvNeXt-S EVA02-B MAE-B CLIP-B32 SigLIP-B
Transfer Transfer Transfer Transfer Transfer Transfer Transfer Transfer Transfer
770 770 770 4k 2k 2k 2k 2k 2k
21.6M 21.7M 22.1M 23.5M 49.5M 85.8M 85.8M 87.5M 92.9M
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
C.2. MLP
Δ Accuracy
Accuracy (%)
Training year
0
1955 −10
1928
−20
1915
Evaluation year
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
19 15
1905
20 2009 13
20 00
19 91
50
Evaluation year
(a) MLP-S
(b) MLP-M
2013 2009
2013 2009
2000
2000
90 1991
20
1991 1982
1955
70
Training year
1964
Accuracy (%)
80
1973
1946
10
1973 1964
0
1955 1946
1937
Δ Accuracy
1982
−10
1937
60
1928
1928
1915
−20
1915
50
Evaluation year
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
19 15
1905
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1905
19 05
1964
1937
60
Evaluation year
Training year
10
1973
1946
19 82
20 2009 13
20 00
19 91
19 82
19 73
19 64
1905 19 55
1915
1905
Evaluation year
70
19 73
−20
1915
19 05
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1928
19 46
50
1905
1937
1928
19 37
1915
1937
19 28
60
1928
19 15
1937
1955 1946
−10
19 64
1946
1964
19 55
1955
80
1973
19 46
0
19 37
1964
Training year
70
1946
1973
19 05
1955
1982
10 Δ Accuracy
1964
Accuracy (%)
80
1973
20
1991
1982
1982 Training year
Training year
1991
1991
1982
2000
90
20
19 28
1991
2013 2009
2000
2000
90
19 05
2013 2009
2013 2009
2000
19 15
2013 2009
Evaluation year
(c) MLP-L Figure 6: MLP models: Accuracy drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
C.3. CNN
1937
Evaluation year
Δ Accuracy
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 05
20 2009 13
20 00
19 91
19 82
19 73
1905
Evaluation year
Evaluation year
(a) CNN-S
(b) CNN-M
2013 2009
2013 2009
2000
2000
90 1991
20
1991 1982
1955
70
Training year
1964
Accuracy (%)
80
1973
1946
10
1973 1964
0
1955 1946
1937
Δ Accuracy
1982
−10
1937
60
1928
1928
1915
−20
1915
50
Evaluation year
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
19 15
1905
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1905
19 05
−20
1915
50
Evaluation year
Training year
−10
1937 1928
19 64
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
1905 19 37
1915
1905
19 28
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1915
19 05
50
1905
19 15
1915
60
1928
−20
19 55
1928
0
1955
19 37
−10
1937
1964
1946
19 46
60
1928
70
1946
19 37
1937
1955
10
1973
19 28
1946
1964
19 15
1955
80
1973
Training year
0
Accuracy (%)
1964
Training year
70
1946
1973
19 05
1955
1982
10 Δ Accuracy
1964
Accuracy (%)
80
1973
20
1991
1982
1982 Training year
Training year
1991
1991
1982
2000
90
20
19 28
1991
2013 2009
2000
2000
90
19 05
2013 2009
2013 2009
2000
19 15
2013 2009
Evaluation year
(c) CNN-L Figure 7: CNN models: Accuracy drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
15
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
C.4. ResNet
Δ Accuracy
Accuracy (%)
Training year
0
1955 −10
1928
−20
1915
Evaluation year
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
19 15
1905
20 2009 13
20 00
19 91
50
Evaluation year
(a) ResNet-S
(b) ResNet-M
2013 2009
2013 2009
2000
2000
90 1991
20
1991 1982
1955
70
Training year
1964
Accuracy (%)
80
1973
1946
10
1973 1964
0
1955 1946
1937
Δ Accuracy
1982
−10
1937
60
1928
1928
1915
−20
1915
50
Evaluation year
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
19 15
1905
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1905
19 05
1964
1937
60
Evaluation year
Training year
10
1973
1946
19 82
20 2009 13
20 00
19 91
19 82
19 73
19 64
1905 19 55
1915
1905
Evaluation year
70
19 73
−20
1915
19 05
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1928
19 46
50
1905
1937
1928
19 37
1915
1937
19 28
60
1928
19 15
1937
1955 1946
−10
19 64
1946
1964
19 55
1955
80
1973
19 46
0
19 37
1964
Training year
70
1946
1973
19 05
1955
1982
10 Δ Accuracy
1964
Accuracy (%)
80
1973
20
1991
1982
1982 Training year
Training year
1991
1991
1982
2000
90
20
19 28
1991
2013 2009
2000
2000
90
19 05
2013 2009
2013 2009
2000
19 15
2013 2009
Evaluation year
(c) ResNet-L Figure 8: ResNet models: Accuracy drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
C.5. ViT
1937
Evaluation year
Δ Accuracy
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 05
20 2009 13
20 00
19 91
19 82
19 73
1905
Evaluation year
Evaluation year
(a) ViT-S
(b) ViT-M
2013 2009
2013 2009
2000
2000
90 1991
20
1991 1982
1955
70
Training year
1964
Accuracy (%)
80
1973
1946
10
1973 1964
0
1955 1946
1937
Δ Accuracy
1982
−10
1937
60
1928
1928
1915
−20
1915
50
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
Evaluation year
19 15
1905
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1905
19 05
−20
1915
50
Evaluation year
Training year
−10
1937 1928
19 64
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
1905 19 37
1915
1905
19 28
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1915
19 05
50
1905
19 15
1915
60
1928
−20
19 55
1928
0
1955
19 37
−10
1937
1964
1946
19 46
60
1928
70
1946
19 37
1937
1955
10
1973
19 28
1946
1964
19 15
1955
80
1973
Training year
0
Accuracy (%)
1964
Training year
70
1946
1973
19 05
1955
1982
10 Δ Accuracy
1964
Accuracy (%)
80
1973
20
1991
1982
1982 Training year
Training year
1991
1991
1982
2000
90
20
19 28
1991
2013 2009
2000
2000
90
19 05
2013 2009
2013 2009
2000
19 15
2013 2009
Evaluation year
(c) ViT-L Figure 9: ViT models: Accuracy drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
16
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
C.6. Transfer
Evaluation year
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
19 05
Δ Accuracy
20 2009 13
Δ Accuracy
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
19 05
Δ Accuracy
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 05
19 28
Δ Accuracy
−20
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 05
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
1905
Evaluation year
Evaluation year
(g) ResNet50-IN
(h) SigLIP-B
2013 2009
2013 2009
2000
2000
90 1991
20
1991 1982
1955
70
Training year
1964
Accuracy (%)
80
1973
1946
10
1973 1964
0
1955 1946
1937
Δ Accuracy
1982
−10
1937
60
1928
1928
1915
−20
1915
50
Evaluation year
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 05
19 15
1905
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1905
19 05
−10
1937
1915
50
Evaluation year
Training year
0
1955
1928
19 15
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
1905 19 28
1915
Evaluation year
60
1928
1905
1964
1946
1937 −20
19 15
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
70
10
1973
19 28
−10
1915
19 05
50
1905
1955 1946
1937 1928
1915
1964
19 15
1955
80
1973
Accuracy (%)
0
19 05
60
1928
1982 Training year
1964
Training year
1973
20
1991
10
1946
1937
Accuracy (%)
1991
Δ Accuracy
70
1946
Accuracy (%)
1955
Training year
Training year
80
1964
2000
90
1982
1982
1973
2013 2009
2000 20
1991
1982
19 05
(f) MAE-B
2000
1991
Training year
20 2009 13
20 00
19 91
19 73
19 64
19 55
19 46
19 37
19 82
Evaluation year
2013 2009
2013 2009
90
1905
Evaluation year
(e) EVA02-B 2000
−20
1915
50
Evaluation year
2013 2009
−10
1928
19 28
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
1905
19 05
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1915
1905
0
1955
1937
60
1928
1915
1964
1946
1937 −20
19 15
50
1905
70
10
1973
19 15
−10
1928
1915
1955 1946
1937
60
1928
1964
Accuracy (%)
1955
80
1973
19 15
1937
1982 Training year
0
1946
Training year
1964
19 05
70
1946
1973
20
1991
10 Δ Accuracy
1955
Accuracy (%)
80
1964
Training year
Training year
1991 1982
1982
1973
2000
90
20
1991
1982
2013 2009
2000
2000
1991
19 05
(d) DINOv3-S 2013 2009
2013 2009
90
Accuracy (%)
20 2009 13
20 00
19 91
19 73
19 64
19 82
Evaluation year
(c) DINOv2-S 2000
−20
1905
Evaluation year
Evaluation year
2013 2009
−10
1915
50
19 55
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
1905 19 46
1915
1905
0
1955
1928
19 46
−20
1915
1964
1937
60
19 37
1928
10
1973
1946
19 28
1937
1928
Evaluation year
70
Training year
1955
1937
19 05
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1964
1946
−10
19 37
50
1905
Training year
1955
80
1973
19 15
0
19 28
1915
19 05
1964
19 15
60
1928
1982
10
1973
20
1991
1982
1946
1937
2000
1991
19 05
1946
2013 2009
90
Δ Accuracy
70
Accuracy (%)
1955
Training year
Training year
80
1964
20 2009 13
20 00
19 91
Evaluation year
2000 20
1982
1973
1905
(b) ConvNeXt-S
1991
1982
−20
1915
50
2013 2009
2000
1991
−10
Evaluation year
2013 2009
90
0
1955
1928
(a) CLIP-B32 2000
1964
1937
60
Evaluation year
2013 2009
10
1973
1946
19 82
20 2009 13
20 00
19 91
19 82
19 73
19 64
1905 19 55
1915
1905
Evaluation year
70
19 73
−20
1915
19 05
20 2009 13
20 00
19 91
19 82
19 73
19 64
19 55
19 46
19 37
19 28
19 15
1928
19 46
50
1905
1937
1928
19 37
1915
1937
19 28
60
1928
19 15
1937
1955 1946
−10
19 64
1946
1964
19 55
1955
80
1973
19 46
0
19 37
1964
Training year
70
1946
1973
19 05
1955
1982
10 Δ Accuracy
1964
Accuracy (%)
80
1973
20
1991
1982
1982 Training year
Training year
1991
1991
1982
2000
90
20
19 28
1991
2013 2009
2000
2000
90
19 05
2013 2009
2013 2009
2000
19 15
2013 2009
Evaluation year
(i) ViT-S16-IN21k Figure 10: Transfer models: Accuracy drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
17
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
C.7. Forgetting and Rankings To see how quickly each model forgets, we summarize its drift matrix as a forgetting curve. The curve plots the Accuracy against the lag ℓ = j − i, the number of slices between the training cutoff i and the evaluation slice j. At each lag we average over all training cutoffs, F (ℓ) = meani Mi, i+ℓ . The result is the Accuracy at a fixed temporal distance, independent of which period a model was trained on. This separates the effect of temporal distance from the difficulty of any single slice.
90 CLIP-B32
90
CNN-L CNN-M CNN-S
70
60
50 0
20
40 60 Years since training
80
Accuracy (%)
Accuracy (%)
80
ConvNeXt-S DINOv2-S DINOv3-S EVA02-B MAE-B MLP-L MLP-M MLP-S ResNet50-IN ResNet-L ResNet-M ResNet-S SigLIP-B ViT-L ViT-M ViT-S ViT-S16-IN21k
80
CNN MLP ResNet Transfer ViT
70
60
50 0
100
20
40 60 Years since training
(a) Per model
80
100
(b) Per model family
Figure 11: Forgetting curves: each model (left) and averaged within each family (right).
Future performance
Decay
CNN-L
79.0
MLP-S
CNN-M
78.4
DINOv3-S
7.9 8.0
7.1
ResNet-L
76.1
ViT-L
ResNet-M
75.8
MAE-B
ResNet-S
75.0
ResNet50-IN
CNN-S
74.7
CLIP-B32
9.1
8.4 8.7
MLP-M
73.0
SigLIP-B
9.2
MLP-L
72.7
ConvNeXt-S
9.2
EVA02-B
9.3
ConvNeXt-S
70.6
ViT-M
ViT-S
66.1
9.6
ViT-S
64.6
DINOv2-S
CLIP-B32
64.1
ViT-M
EVA02-B
63.4
ViT-S16-IN21k
DINOv2-S
62.7
MLP-L
12.2
SigLIP-B
62.4
ResNet-S
12.3
MLP-S
62.3
CNN-S
12.3
ResNet50-IN
61.9
MLP-M
ViT-S16-IN21k
61.8
ResNet-M
ViT-L
60.3
MAE-B
59.0
DINOv3-S
57.9
0
10
20
30
40 50 Accuracy (%)
60
70
80
9.8 10.1 10.4
12.7 13.1
CNN-L
13.7
ResNet-L
13.9
CNN-M
13.9
0
(a) By future performance
2
4
6
8 Decay (%)
(b) By decay
Figure 12: Models ranked by mean future performance and by temporal decay.
18
10
12
14
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
C.8. Result Tables
Table 5: Temporal robustness on Yearbook. Model MLP-S DINOv3-S ViT-L MAE-B ResNet50-IN CLIP-B32 SigLIP-B ConvNeXt-S EVA02-B ViT-S DINOv2-S
Future (%)
Decay (%)
62.3 57.9 60.3 59.0 61.9 64.1 62.4 70.6 63.4 64.6 62.7
7.1 7.9 8.0 8.4 8.7 9.1 9.2 9.2 9.3 9.6 9.8
Model ViT-M ViT-S16-IN21k MLP-L ResNet-S CNN-S MLP-M ResNet-M CNN-L ResNet-L CNN-M
Future (%)
Decay (%)
66.1 61.8 72.7 75.0 74.7 73.0 75.8 79.0 76.1 78.4
10.1 10.4 12.2 12.3 12.3 12.7 13.1 13.7 13.9 13.9
Table 6: Yearbook: models trained up to 1905, ordered by Table 7: Yearbook: models trained up to 1944, ordered by future performance. future performance. Rank
Model
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
ViT-S16-IN21k ResNet-M ResNet-L MLP-L MAE-B ResNet-S DINOv2-S DINOv3-S ViT-L MLP-S ConvNeXt-S ViT-M CNN-L ViT-S ResNet50-IN EVA02-B MLP-M CNN-S CLIP-B32 CNN-M SigLIP-B
Accuracy (%)
Future (%)
Decay (%)
Rank
Model
50.0 60.0 70.0 40.0 40.0 70.0 50.0 50.0 50.0 40.0 50.0 50.0 50.0 50.0 50.0 50.0 50.0 50.0 50.0 60.0 50.0
50.7 50.3 50.1 49.8 49.5 49.3 48.5 48.3 46.8 46.3 46.1 45.7 45.4 45.1 45.1 44.9 44.9 44.9 44.7 44.7 44.7
-0.7 9.7 19.9 -9.8 -9.5 20.7 1.5 1.7 3.2 -6.3 3.9 4.3 4.6 4.9 4.9 5.1 5.1 5.1 5.3 15.3 5.3
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
ResNet-L ResNet-M CNN-M ResNet-S CNN-L CNN-S ConvNeXt-S MLP-M MLP-L ViT-M ViT-L ViT-S CLIP-B32 EVA02-B DINOv2-S SigLIP-B ViT-S16-IN21k ResNet50-IN MAE-B DINOv3-S MLP-S
19
Accuracy (%)
Future (%)
Decay (%)
98.4 98.3 98.8 98.0 97.1 95.0 96.3 93.4 87.9 92.6 91.0 89.4 94.3 93.1 88.5 88.9 86.4 80.2 86.7 82.7 65.8
82.9 81.9 81.4 81.1 80.2 79.1 78.5 77.6 74.6 71.4 70.7 69.9 69.9 67.9 67.1 66.8 65.0 64.8 62.6 62.4 57.9
15.5 16.4 17.4 16.9 16.8 15.9 17.8 15.9 13.3 21.2 20.3 19.5 24.5 25.2 21.5 22.1 21.4 15.3 24.1 20.3 7.9
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Table 8: Yearbook: models trained up to 1978, ordered by Table 9: Yearbook: models trained up to 2012, ordered by future performance. future performance. Rank
Model
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
CNN-L ResNet-M ResNet-L CNN-M MLP-M ResNet-S MLP-L CNN-S ConvNeXt-S ViT-M CLIP-B32 EVA02-B ViT-L DINOv2-S ViT-S ViT-S16-IN21k ResNet50-IN SigLIP-B MLP-S MAE-B DINOv3-S
Accuracy (%)
Future (%)
Decay (%)
Rank
Model
85.9 88.7 87.6 86.8 79.4 83.4 75.5 72.7 80.6 72.4 69.9 60.6 66.8 64.5 64.2 66.5 72.7 64.8 67.9 56.1 60.3
91.4 90.4 89.3 88.8 84.5 84.4 83.2 81.1 78.6 76.8 74.8 73.3 72.4 72.3 71.8 70.7 70.7 69.9 69.4 65.0 59.9
-5.4 -1.6 -1.7 -2.0 -5.1 -1.0 -7.7 -8.5 1.9 -4.4 -5.0 -12.8 -5.6 -7.8 -7.6 -4.3 2.0 -5.1 -1.5 -8.9 0.4
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
CNN-L CNN-M ResNet-L ResNet-M ResNet-S CNN-S MLP-L MLP-M ConvNeXt-S DINOv2-S ViT-M ViT-S SigLIP-B EVA02-B CLIP-B32 ViT-S16-IN21k ResNet50-IN DINOv3-S MAE-B MLP-S ViT-L
Accuracy (%)
Future (%)
Decay (%)
93.2 93.2 88.9 92.6 86.8 83.7 87.4 86.3 73.7 82.6 77.9 80.5 85.8 82.1 74.7 87.9 76.8 71.1 78.4 70.0 66.8
98.2 97.7 97.5 97.0 95.4 93.4 91.9 91.7 91.2 90.0 88.1 88.0 87.5 87.4 86.6 86.5 83.8 77.9 77.4 76.1 74.4
-5.0 -4.6 -8.5 -4.3 -8.5 -9.7 -4.6 -5.4 -17.5 -7.3 -10.2 -7.5 -1.7 -5.3 -11.9 1.4 -7.0 -6.9 1.0 -6.1 -7.5
Table 10: Yearbook: future performance and decay by model family.
Family CNN MLP ResNet Transfer ViT
1905 Future Decay 45.0 47.0 49.9 46.9 45.9
8.3 -3.7 16.8 1.9 4.1
1944 Future Decay 80.3 70.0 82.0 67.2 70.7
16.7 12.4 16.2 21.4 20.3
20
1978 Future Decay 87.1 79.0 88.0 70.6 73.7
-5.3 -4.8 -1.4 -4.4 -5.9
2012 Future Decay 96.4 86.6 96.6 85.4 83.5
-6.4 -5.3 -7.1 -6.1 -8.4
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
D. Amazon Reviews – Drift Matrices The cohort-mean and per-model deviation matrices shown here, and the in-distribution, future, and decay quantities tabulated below, are defined in Section 4.5.
2023-H2
1.05
2023-H1 1.00
2022-H1
Training half-year
2020-H1
0.90
2019-H1
0.85
2018-H1
0.80
2017-H1
Balanced MSE
0.95
2021-H1
0.75
2016-H1 0.70 2015-H1 0.65
H 1 20 23 20 -H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
20 14 -
H 1
2014-H1
Evaluation half-year
Figure 13: Cohort-mean Balanced MSE matrix M̄ over the Amazon Reviews models. Cell (i, j) is the mean across those models of the score from training through slice i and evaluating on slice j.
D.1. Model Roster Table 11: Amazon Reviews: models trained from scratch. Model
Family
TX-S TextCNN-S FFN-S BiGRU-S FFN-M TextCNN-M TX-M BiLSTM-M FFN-L TX-L TextCNN-L BiLSTM-Attn-L
Transformer TextCNN FFN Recurrent FFN TextCNN Transformer Recurrent FFN Transformer TextCNN Recurrent
Trainable
Total
83k 88k 99k 99k 394k 461k 493k 536k 1.6M 1.9M 1.9M 2.2M
124.7M 124.7M 124.7M 124.7M 125.0M 125.1M 125.1M 125.2M 126.2M 126.5M 126.6M 126.8M
Table 12: Amazon Reviews: frozen pretrained encoders, with trainable head and total parameters.
21
Model
Family
Trainable
Total
DistilBERT ELECTRA BERT MPNet RoBERTa ModernBERT DeBERTa-v3
Frozen Frozen Frozen Frozen Frozen Frozen Frozen
769 769 769 769 769 769 769
66.4M 108.9M 109.5M 109.5M 124.6M 149.0M 183.8M
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
D.2. FFN
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
2020-H1
0.90
2019-H1
0.85
2018-H1
0.80
2017-H1
−0.1
H 1
H 1
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
20 16 -
20 14 -
Evaluation half-year
Evaluation half-year
Evaluation half-year
(a) FFN-S
(b) FFN-M
2023-H2
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
0.85
2018-H1
0.80
2017-H1
Training half-year
0.90
2019-H1
2021-H1
Balanced MSE
2020-H1
0.2
2022-H1
0.95
2021-H1 Training half-year
20 15 -
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
20 16 -
20 15 -
H 1
2014-H1
H 1 20 14 -
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
H 1
H 1
H 1
H 1
Evaluation half-year
20 19 -
20 18 -
20 17 -
20 16 -
20 14 -
20 15 -
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
H 1
H 1
H 1
H 1
H 1
−0.2
0.65
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
−0.1
2015-H1
2014-H1
2014-H1
20 16 -
0.0
2018-H1
2016-H1
2015-H1
−0.2
0.65 2014-H1
20 15 -
2019-H1
0.70 2015-H1
H 1
0.1
2020-H1
2017-H1
0.75
0.70 2015-H1
20 14 -
2021-H1
2016-H1
2016-H1
0.2
2022-H1
0.95
Δ Balanced MSE
2021-H1
Balanced MSE
0.80
2017-H1
2021-H1
1.00
Training half-year
2018-H1
2022-H1
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
Δ Balanced MSE
0.85
2023-H1
0.2
2022-H1
Training half-year
0.90
2019-H1
Balanced MSE
2020-H1
Training half-year
0.95
2023-H2
1.05
2023-H1
2023-H1
1.00
2021-H1 Training half-year
2023-H2
2023-H2
1.05
2022-H1
Δ Balanced MSE
2023-H2 2023-H1
−0.1
2016-H1
0.70 2015-H1
2015-H1
−0.2
0.65 2014-H1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
H 1
20 20 -
H 1
Evaluation half-year
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
20 14 -
H 1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
20 14 -
H 1
2014-H1
Evaluation half-year
(c) FFN-L Figure 14: FFN models: Balanced MSE drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
D.3. TextCNN
2017-H1
0.85
2018-H1
0.80
2017-H1
−0.1
0.75
2016-H1
2016-H1 2015-H1
2015-H1
−0.2
−0.1
2015-H1
−0.2
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
Evaluation half-year
Evaluation half-year
Evaluation half-year
(a) TextCNN-S
(b) TextCNN-M
2023-H2
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
0.85
2018-H1
0.80
2017-H1
Training half-year
0.90
2019-H1
2021-H1
Balanced MSE
2020-H1
0.2
2022-H1
0.95
2021-H1 Training half-year
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
20 16 -
20 15 -
H 1
2014-H1
H 1 20 14 -
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
Evaluation half-year
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 14 -
20 15 -
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
2017-H1
0.65 2014-H1
2014-H1
H 1 20 14 -
0.0
2018-H1
0.70
0.65 2014-H1
2019-H1
2016-H1
0.70 2015-H1
0.1
2020-H1
Δ Balanced MSE
0.90
2019-H1
20 14 -
0.75
2016-H1
0.0
2018-H1
2020-H1
Training half-year
0.80
2017-H1
2019-H1
Balanced MSE
2018-H1
2020-H1
2021-H1
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
Δ Balanced MSE
0.85
0.2
2022-H1
0.95
2021-H1
0.1
Training half-year
0.90
2019-H1
Balanced MSE
2020-H1
1.00
2022-H1
2021-H1 Training half-year
2021-H1
2023-H1
0.2
2022-H1
0.95
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
Training half-year
2023-H2
2023-H2
1.05
2023-H1
Δ Balanced MSE
2023-H2
−0.1
2016-H1
0.70 2015-H1
2015-H1
−0.2
0.65 2014-H1
Evaluation half-year
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
20 14 -
H 1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
20 14 -
H 1
2014-H1
Evaluation half-year
(c) TextCNN-L Figure 15: TextCNN models: Balanced MSE drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
22
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
D.4. Recurrent
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
2020-H1
0.90
2019-H1
0.85
2018-H1
0.80
2017-H1
−0.1
H 1
H 1
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
20 16 -
20 14 -
Evaluation half-year
Evaluation half-year
Evaluation half-year
(a) BiGRU-S
(b) BiLSTM-M
2023-H2
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
0.85
2018-H1
0.80
2017-H1
Training half-year
0.90
2019-H1
2021-H1
Balanced MSE
2020-H1
0.2
2022-H1
0.95
2021-H1 Training half-year
20 15 -
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
20 16 -
20 15 -
H 1
2014-H1
H 1 20 14 -
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
H 1
H 1
H 1
H 1
Evaluation half-year
20 19 -
20 18 -
20 17 -
20 16 -
20 14 -
20 15 -
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
H 1
H 1
H 1
H 1
H 1
−0.2
0.65
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
−0.1
2015-H1
2014-H1
2014-H1
20 16 -
0.0
2018-H1
2016-H1
2015-H1
−0.2
0.65 2014-H1
20 15 -
2019-H1
0.70 2015-H1
H 1
0.1
2020-H1
2017-H1
0.75
0.70 2015-H1
20 14 -
2021-H1
2016-H1
2016-H1
0.2
2022-H1
0.95
Δ Balanced MSE
2021-H1
Balanced MSE
0.80
2017-H1
2021-H1
1.00
Training half-year
2018-H1
2022-H1
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
Δ Balanced MSE
0.85
2023-H1
0.2
2022-H1
Training half-year
0.90
2019-H1
Balanced MSE
2020-H1
Training half-year
0.95
2023-H2
1.05
2023-H1
2023-H1
1.00
2021-H1 Training half-year
2023-H2
2023-H2
1.05
2022-H1
Δ Balanced MSE
2023-H2 2023-H1
−0.1
2016-H1
0.70 2015-H1
2015-H1
−0.2
0.65 2014-H1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
H 1
20 20 -
H 1
Evaluation half-year
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
20 14 -
H 1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
20 14 -
H 1
2014-H1
Evaluation half-year
(c) BiLSTM-Attn-L Figure 16: Recurrent models: Balanced MSE drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
D.5. Transformer
2017-H1
0.85
2018-H1
0.80
2017-H1
−0.1
0.75
2016-H1
2016-H1 2015-H1
2015-H1
−0.2
−0.1
2015-H1
−0.2
Evaluation half-year
Evaluation half-year
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
Evaluation half-year
(a) TX-S
(b) TX-M
2023-H2
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
0.85
2018-H1
0.80
2017-H1
Training half-year
0.90
2019-H1
2021-H1
Balanced MSE
2020-H1
0.2
2022-H1
0.95
2021-H1 Training half-year
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
20 16 -
20 15 -
H 1
2014-H1
H 1 20 14 -
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
Evaluation half-year
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 14 -
20 15 -
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
2017-H1
0.65 2014-H1
2014-H1
H 1 20 14 -
0.0
2018-H1
0.70
0.65 2014-H1
2019-H1
2016-H1
0.70 2015-H1
0.1
2020-H1
Δ Balanced MSE
0.90
2019-H1
20 14 -
0.75
2016-H1
0.0
2018-H1
2020-H1
Training half-year
0.80
2017-H1
2019-H1
Balanced MSE
2018-H1
2020-H1
2021-H1
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
Δ Balanced MSE
0.85
0.2
2022-H1
0.95
2021-H1
0.1
Training half-year
0.90
2019-H1
Balanced MSE
2020-H1
1.00
2022-H1
2021-H1 Training half-year
2021-H1
2023-H1
0.2
2022-H1
0.95
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
Training half-year
2023-H2
2023-H2
1.05
2023-H1
Δ Balanced MSE
2023-H2
−0.1
2016-H1
0.70 2015-H1
2015-H1
−0.2
0.65 2014-H1
Evaluation half-year
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
20 14 -
H 1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
20 14 -
H 1
2014-H1
Evaluation half-year
(c) TX-L Figure 17: Transformer models: Balanced MSE drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
23
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
D.6. Frozen
2017-H1
0.75
2016-H1
0.80
2017-H1
−0.1
0.75
2015-H1
−0.2
2015-H1
0.65
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
0.90
2019-H1
0.85
2018-H1
0.80
2017-H1
−0.1
1.00
2023-H1
2019-H1
0.0
2018-H1 2017-H1
Training half-year
2020-H1
2020-H1
0.90
2019-H1
0.85
2018-H1
0.80
2017-H1
−0.1
0.75
2016-H1
2016-H1 2015-H1
2015-H1
−0.2
H 1
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
2017-H1
−0.1
2015-H1
−0.2
Evaluation half-year
Evaluation half-year
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
Evaluation half-year
(e) ModernBERT
(f) RoBERTa
2023-H2
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
0.85
2018-H1
0.80
2017-H1
Training half-year
0.90
2019-H1
2021-H1
Balanced MSE
2020-H1
0.2
2022-H1
0.95
2021-H1 Training half-year
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
20 16 -
20 15 -
H 1
2014-H1
H 1 20 14 -
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
Evaluation half-year
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 14 -
20 15 -
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
0.0
2018-H1
0.65 2014-H1
2014-H1
H 1 20 14 -
2019-H1
0.70
0.65 2014-H1
0.1
2020-H1
2016-H1
0.70 2015-H1
2021-H1
20 15 -
2016-H1
2021-H1
0.2
2022-H1
0.95
20 14 -
0.75
2023-H1
1.00
2022-H1
0.1
2023-H2
1.05
0.2
Δ Balanced MSE
Balanced MSE
0.80
2017-H1
2023-H2
2023-H1
2021-H1 Training half-year
Training half-year
2018-H1
(d) MPNet
2023-H2
2022-H1
0.95
20 14 -
Evaluation half-year
Δ Balanced MSE
1.05
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
H 1 20 14 -
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 14 -
20 15 -
H 1
H 1 20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
H 1 20 14 -
20 21 -
Evaluation half-year
(c) ELECTRA 2023-H1
0.85
−0.2
2014-H1
Evaluation half-year
0.90
−0.1
2015-H1
0.65
Evaluation half-year
2019-H1
0.0
2018-H1
2016-H1
2014-H1
2014-H1
2020-H1
2019-H1
2017-H1
0.75
2015-H1
−0.2
0.65 2014-H1
2021-H1
0.1
2020-H1
0.70 2015-H1
2022-H1
2021-H1
2016-H1
2016-H1
0.70 2015-H1
2023-H2
H 1
2020-H1
0.2
2022-H1
0.95
Δ Balanced MSE
2021-H1
Balanced MSE
2021-H1
1.00
Training half-year
0.80
2017-H1
2022-H1
Balanced MSE
2018-H1
2022-H1
Training half-year
0.85
2023-H1
0.2
Δ Balanced MSE
0.90
2019-H1
Balanced MSE
0.95
2023-H2
1.05
2023-H1
2023-H1
Training half-year
Training half-year
(b) DistilBERT 2023-H2
2023-H2
1.00
20 16 -
Evaluation half-year
(a) BERT 1.05
20 14 -
Evaluation half-year
Evaluation half-year
2023-H1
20 15 -
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
20 16 -
20 15 -
H 1
2014-H1
H 1 20 14 -
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
H 1
H 1
H 1
H 1
Evaluation half-year
20 19 -
20 18 -
20 17 -
20 16 -
20 14 -
20 15 -
H 1
H 1
H 1
H 1
H 1
H 1
H 1
H 1 20 2 20 3-H 23 1 -H 2
20 22 -
20 21 -
20 20 -
20 19 -
20 18 -
20 17 -
20 16 -
20 15 -
H 1 20 14 -
2014-H1
2014-H1
2020-H1
−0.2
0.65
2014-H1
2021-H1
−0.1
0.70 2015-H1
2022-H1
0.0
2018-H1
2016-H1
0.70 2015-H1
2023-H2
2019-H1
2017-H1
2016-H1
2016-H1
0.1
2020-H1
Δ Balanced MSE
0.85
2018-H1
Balanced MSE
0.0
2018-H1
Training half-year
2019-H1
0.90
2019-H1
Training half-year
0.80
2017-H1
2020-H1
2021-H1
0.1
2020-H1 2019-H1
0.0
2018-H1 2017-H1
0.75
2016-H1
Δ Balanced MSE
0.85
2018-H1
2020-H1
0.2
2022-H1
0.95
2021-H1
0.1
Training half-year
2019-H1
Balanced MSE
0.90
2020-H1
1.00
2022-H1
2021-H1 Training half-year
2021-H1
2023-H1
0.2
2022-H1
0.95
2023-H2
1.05
2023-H1
2023-H1
1.00
2022-H1
Training half-year
2023-H2
2023-H2
1.05
2023-H1
Δ Balanced MSE
2023-H2
−0.1
2016-H1
0.70 2015-H1
2015-H1
−0.2
0.65 2014-H1
Evaluation half-year
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
20 21 -
H 1
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
20 17 -
H 1
H 1
20 16 -
20 15 -
20 14 -
H 1
20 2 20 3-H 23 1 -H 2
H 1
20 22 -
H 1
H 1
20 21 -
H 1
20 20 -
H 1
20 19 -
H 1
20 18 -
H 1
20 17 -
H 1
20 16 -
20 15 -
20 14 -
H 1
2014-H1
Evaluation half-year
(g) DeBERTa-v3 Figure 18: Frozen models: Balanced MSE drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
D.7. Forgetting and Rankings To see how quickly each model forgets, we summarize its drift matrix as a forgetting curve. The curve plots the Balanced MSE against the lag ℓ = j − i, the number of slices between the training cutoff i and the evaluation slice j. At each lag we average over all training cutoffs, F (ℓ) = meani Mi, i+ℓ .
The result is the Balanced MSE at a fixed temporal distance, independent of which period a model was trained on. This separates the effect of temporal distance from the difficulty of any single slice. 24
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift 1.2 BERT BiGRU-S BiLSTM-Attn-L BiLSTM-M DeBERTa-v3
1.2
1.1
DistilBERT ELECTRA FFN-L FFN-M FFN-S ModernBERT MPNet RoBERTa TextCNN-L
1.0
0.9
Balanced MSE
Balanced MSE
1.1
TextCNN-M TextCNN-S TX-L TX-M
0.8
1.0
FFN Frozen Recurrent TextCNN Transformer
0.9
0.8
TX-S
0.7
0.7 0.0
2.5
5.0
7.5 10.0 12.5 Half-years since training
15.0
0.0
17.5
2.5
5.0
(a) Per model
7.5 10.0 12.5 Half-years since training
15.0
17.5
(b) Per model family
Figure 19: Forgetting curves: each model (left) and averaged within each family (right).
Future performance TX-S
Decay DeBERTa-v3
0.776
0.043
BiGRU-S
0.797
ModernBERT
TX-M
0.801
BERT
0.075
0.064
BiLSTM-Attn-L
0.819
FFN-S
0.075
BiLSTM-M
0.824
ELECTRA
0.076
TX-L
0.829
FFN-M
0.077
FFN-M
0.834
TextCNN-L
0.078
FFN-S
0.841
DistilBERT
FFN-L
0.848
RoBERTa
0.088
DeBERTa-v3
0.849
FFN-L
0.089
TextCNN-S
0.860
TX-S
0.090
TextCNN-L
0.871
TX-M
0.092
TextCNN-M
0.885
TextCNN-M
0.093
TX-L
0.094
RoBERTa
0.944
BERT
0.999
MPNet
ELECTRA
1.000
BiLSTM-Attn-L
MPNet
1.047
TextCNN-S
DistilBERT
1.047
BiGRU-S
ModernBERT 0.0
0.4
0.6 Balanced MSE
0.8
0.099 0.106 0.114 0.116
BiLSTM-M
1.086
0.2
0.082
1.0
0.00
0.128
0.02
(a) By future performance
0.04
0.06 0.08 Decay
0.10
(b) By decay
Figure 20: Models ranked by mean future performance and by temporal decay.
D.8. Result Tables Table 13: Temporal robustness on Amazon Reviews. Model
Future
Decay
DeBERTa-v3 ModernBERT BERT FFN-S ELECTRA FFN-M TextCNN-L DistilBERT RoBERTa FFN-L
0.849 1.086 0.999 0.841 1.000 0.834 0.871 1.047 0.944 0.848
0.043 0.064 0.075 0.075 0.076 0.077 0.078 0.082 0.088 0.089
25
Model
Future
Decay
TX-S TX-M TextCNN-M TX-L MPNet BiLSTM-Attn-L TextCNN-S BiGRU-S BiLSTM-M
0.776 0.801 0.885 0.829 1.047 0.819 0.860 0.797 0.824
0.090 0.092 0.093 0.094 0.099 0.106 0.114 0.116 0.128
0.12
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Table 14: Amazon Reviews: models trained up to 2014-H1, Table 15: Amazon Reviews: models trained up to 2017-H1, ordered by future performance. ordered by future performance. Rank
Model
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
TX-M TextCNN-L FFN-L TX-L FFN-M TX-S TextCNN-M FFN-S BiGRU-S TextCNN-S DeBERTa-v3 BiLSTM-Attn-L BiLSTM-M BERT DistilBERT ELECTRA ModernBERT RoBERTa MPNet
Balanced MSE
Future
Decay
Rank
Model
0.748 0.782 0.836 0.767 0.885 0.868 0.761 0.907 0.790 0.782 1.035 0.938 0.850 1.067 1.113 1.332 1.273 1.227 1.492
0.933 0.935 0.972 0.981 0.986 0.987 0.993 1.023 1.073 1.078 1.083 1.105 1.127 1.232 1.250 1.382 1.386 1.468 1.657
0.185 0.153 0.136 0.215 0.100 0.119 0.231 0.116 0.283 0.296 0.048 0.167 0.277 0.165 0.137 0.050 0.113 0.241 0.165
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
TX-S BiGRU-S TX-M BiLSTM-Attn-L FFN-L FFN-S DeBERTa-v3 TextCNN-S TX-L BiLSTM-M FFN-M RoBERTa TextCNN-L TextCNN-M ELECTRA MPNet BERT ModernBERT DistilBERT
Balanced MSE
Future
Decay
0.625 0.644 0.701 0.654 0.724 0.723 0.758 0.711 0.722 0.681 0.742 0.792 0.789 0.773 0.873 0.868 0.929 0.958 0.942
0.748 0.765 0.785 0.789 0.810 0.812 0.814 0.817 0.818 0.828 0.829 0.860 0.876 0.885 0.926 0.940 0.984 1.028 1.047
0.123 0.121 0.084 0.136 0.086 0.089 0.056 0.106 0.096 0.147 0.087 0.068 0.087 0.113 0.052 0.072 0.055 0.070 0.104
Table 16: Amazon Reviews: models trained up to 2020-H1, Table 17: Amazon Reviews: models trained up to 2023-H1, ordered by future performance. ordered by future performance. Rank
Model
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
BiLSTM-M BiGRU-S TX-S BiLSTM-Attn-L TX-M TextCNN-S TX-L FFN-L FFN-S FFN-M DeBERTa-v3 TextCNN-M TextCNN-L RoBERTa ELECTRA MPNet BERT DistilBERT ModernBERT
Balanced MSE
Future
Decay
Rank
Model
0.614 0.627 0.637 0.647 0.675 0.672 0.693 0.713 0.729 0.744 0.753 0.748 0.743 0.800 0.844 0.860 0.876 0.926 0.969
0.706 0.715 0.721 0.749 0.751 0.767 0.773 0.793 0.795 0.817 0.820 0.833 0.834 0.870 0.911 0.938 0.959 0.998 1.041
0.092 0.088 0.084 0.102 0.076 0.094 0.080 0.080 0.066 0.074 0.067 0.085 0.091 0.070 0.067 0.078 0.084 0.072 0.072
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
TX-M BiGRU-S BiLSTM-M BiLSTM-Attn-L TX-S TextCNN-S TX-L DeBERTa-v3 FFN-L FFN-S FFN-M TextCNN-M RoBERTa TextCNN-L ELECTRA BERT MPNet DistilBERT ModernBERT
Balanced MSE
Future
Decay
0.634 0.627 0.638 0.687 0.663 0.682 0.699 0.776 0.728 0.731 0.742 0.751 0.812 0.794 0.882 0.911 0.863 0.937 0.992
0.665 0.691 0.703 0.735 0.740 0.747 0.759 0.764 0.770 0.777 0.821 0.838 0.847 0.869 0.894 0.939 0.956 0.980 1.069
0.031 0.064 0.065 0.048 0.077 0.066 0.060 -0.013 0.042 0.046 0.079 0.087 0.035 0.075 0.012 0.029 0.093 0.043 0.076
Table 18: Amazon Reviews: future performance and decay by model family.
Family
2014-H1 Future Decay
2017-H1 Future Decay
2020-H1 Future Decay
2023-H1 Future Decay
FFN Frozen Recurrent TextCNN Transformer
0.993 1.351 1.101 1.002 0.967
0.817 0.943 0.794 0.859 0.784
0.802 0.934 0.723 0.811 0.748
0.789 0.921 0.710 0.818 0.721
0.117 0.131 0.242 0.227 0.173
0.087 0.068 0.134 0.102 0.101
26
0.073 0.073 0.094 0.090 0.080
0.056 0.039 0.059 0.076 0.056
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
E. arXiv – Drift Matrices The cohort-mean and per-model deviation matrices shown here, and the in-distribution, future, and decay quantities tabulated below, are defined in Section 4.5. The label space comprises the leaf categories cs.LG, hep-ph, cs.CV, cs.AI, hep-th, quant-ph, and gr-qc.
2025 2024
98
2021 96
2015
94
2012 92 2009 2006
Macro AUC (%)
Training year
2018
90
2003 88
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
2000
Evaluation year
Figure 21: Cohort-mean Macro AUC matrix M̄ over the arXiv models. Cell (i, j) is the mean across those models of the score from training through slice i and evaluating on slice j.
E.1. Model Roster Table 19: arXiv: models trained from scratch. Model
Family
TX-S TextCNN-S FFN-S BiGRU-S FFN-M TextCNN-M TX-M BiLSTM-M FFN-L TX-L TextCNN-L BiLSTM-Attn-L
Transformer TextCNN FFN Recurrent FFN TextCNN Transformer Recurrent FFN Transformer TextCNN Recurrent
Trainable
Total
83k 89k 99k 100k 397k 462k 493k 537k 1.6M 1.9M 1.9M 2.2M
124.7M 124.7M 124.7M 124.7M 125.0M 125.1M 125.1M 125.2M 126.2M 126.5M 126.6M 126.8M
Table 20: arXiv: frozen pretrained encoders, with trainable head and total parameters.
27
Model
Family
Trainable
Total
MiniLM-L6 DistilBERT ELECTRA BERT MPNet RoBERTa ModernBERT DeBERTa-v3
Frozen Frozen Frozen Frozen Frozen Frozen Frozen Frozen
3k 5k 5k 5k 5k 5k 5k 5k
22.7M 66.4M 108.9M 109.5M 109.5M 124.7M 149.0M 183.8M
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
E.2. FFN
2025 2024
98
2021
2021
10.0
2025 2024
7.5
2021
90
2003
0.0
2012
94
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
2000
−10.0
2000
90
2.5
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
Evaluation year
−10.0
Evaluation year
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 00
20 03
88
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
5.0
2015
Δ Macro AUC
2015
88 2000
20 00
7.5
2018
Macro AUC (%)
2.5
Training year
2018
20 03
2006
2015
5.0
Δ Macro AUC
92 2009
Macro AUC (%)
2012
Training year
Training year
2018
94
10.0
2021
96
2018 2015
2025 2024
98
96
Training year
2025 2024
Evaluation year
(a) FFN-S
(b) FFN-M
2025 2024
2025 2024
98
2021
10.0
7.5
2021
96 5.0
94
2012 92 2009 2006
90
2003
2.5
2015
Δ Macro AUC
2015
Training year
2018
Macro AUC (%)
Training year
2018
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
88 −10.0
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
2000
Evaluation year
(c) FFN-L Figure 22: FFN models: Macro AUC drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
E.3. TextCNN
2025 2024
98
2021
2021
10.0
2025 2024
7.5
2021
94
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
2000
−10.0
2000
90
Evaluation year
0.0
2009
−2.5
2006
−5.0
2003
−7.5
−10.0
Evaluation year
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
2.5
2012
88 20 03
88 2000
5.0
2015
Δ Macro AUC
0.0
2012
Training year
2015
20 09
2003
2.5
20 06
90
7.5
2018
20 03
2006
2018
Training year
2009
2015
5.0
Δ Macro AUC
92
Macro AUC (%)
2012
Training year
Training year
2018
94
10.0
2021
96
2018 2015
2025 2024
98
96
Macro AUC (%)
2025 2024
Evaluation year
(a) TextCNN-S
(b) TextCNN-M
2025 2024
2025 2024
98
2021
10.0
7.5
2021
96
94
2012 92 2009 2006
Training year
2015
90
2003
2.5
2015
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
Δ Macro AUC
5.0
2018
Macro AUC (%)
Training year
2018
88 −10.0
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
2000
Evaluation year
(c) TextCNN-L Figure 23: TextCNN models: Macro AUC drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
28
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
E.4. Recurrent
2025 2024
98
2021
2021
10.0
2025 2024
7.5
2021
90
2003
0.0
2012
94
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
2000
−10.0
2000
90
2.5
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
Evaluation year
−10.0
Evaluation year
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 00
20 03
88
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
5.0
2015
Δ Macro AUC
2015
88 2000
20 00
7.5
2018
Macro AUC (%)
2.5
Training year
2018
20 03
2006
2015
5.0
Δ Macro AUC
92 2009
Macro AUC (%)
2012
Training year
Training year
2018
94
10.0
2021
96
2018 2015
2025 2024
98
96
Training year
2025 2024
Evaluation year
(a) BiGRU-S
(b) BiLSTM-M
2025 2024
2025 2024
98
2021
10.0
7.5
2021
96 5.0
94
2012 92 2009 2006
90
2003
2.5
2015
Δ Macro AUC
2015
Training year
2018
Macro AUC (%)
Training year
2018
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
88 −10.0
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
2000
Evaluation year
(c) BiLSTM-Attn-L Figure 24: Recurrent models: Macro AUC drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
E.5. Transformer
2025 2024
98
2021
2021
10.0
2025 2024
7.5
2021
94
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
2000
−10.0
2000
90
Evaluation year
0.0
2009
−2.5
2006
−5.0
2003
−7.5
−10.0
Evaluation year
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
2.5
2012
88 20 03
88 2000
5.0
2015
Δ Macro AUC
0.0
2012
Training year
2015
20 09
2003
2.5
20 06
90
7.5
2018
20 03
2006
2018
Training year
2009
2015
5.0
Δ Macro AUC
92
Macro AUC (%)
2012
Training year
Training year
2018
94
10.0
2021
96
2018 2015
2025 2024
98
96
Macro AUC (%)
2025 2024
Evaluation year
(a) TX-S
(b) TX-M
2025 2024
2025 2024
98
2021
10.0
7.5
2021
96
94
2012 92 2009 2006
Training year
2015
90
2003
2.5
2015
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
Δ Macro AUC
5.0
2018
Macro AUC (%)
Training year
2018
88 −10.0
Evaluation year
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
2000
20 03
20 00
2000
Evaluation year
(c) TX-L Figure 25: Transformer models: Macro AUC drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
29
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
E.6. Frozen
2025 2024
98
2021
2021
10.0
2025 2024
7.5
2021
90
2003
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
2000
−10.0
2000
90
2021
10.0
2025 2024
7.5
2021
90
2003
94
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
2000
−10.0
2000
90
10.0
2025 2024
7.5
2021
5.0
2018
2.5
2015
2003
92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
90
2021
2021
10.0
2025 2024
7.5
2021
5.0
2018
2.5
2015
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 2024 25 20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
94
2012 92 2009
2009
−2.5
2006
−5.0
2006
2003
−7.5
2003
90
2000
2.5
0.0
2012 2009
−2.5
2006
−5.0
2003
−7.5
Evaluation year
Evaluation year
Evaluation year
(g) DeBERTa-v3
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
−10.0
2000
20 03
−10.0
20 03
20 00
20 2024 25
5.0
2015
88 2000
20 21
7.5
2018
Δ Macro AUC
Training year
0.0
2012
20 00
2000
20 18
10.0
2021
88
20 15
−10.0
2025 2024
98
20 00
2003
2015
Δ Macro AUC
Macro AUC (%)
Training year
90
20 12
−7.5
96 2018
20 09
−5.0
2003
Evaluation year
96 2018
20 06
−2.5
2006
(f) RoBERTa
2025 2024
98
20 03
0.0
2009
Evaluation year
Evaluation year
(e) ModernBERT
Training year
2.5
2012
2000
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
2000
−10.0
2000
Evaluation year
20 00
5.0
2015
Δ Macro AUC
2012
Macro AUC (%)
94
88
2000
2006
7.5
2018
88
92
10.0
2021
Training year
0.0
2012
Training year
2015
Δ Macro AUC
Macro AUC (%)
Training year
Training year
90
2009
2025 2024
98
96 2018
2012
−10.0
Evaluation year
96 2018
94
−7.5
(d) MPNet
2021
2015
−5.0
2003
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00 2021
2025 2024
−2.5
2006
Evaluation year
2025 2024
98
2006
0.0
2009
2000
(c) ELECTRA
92
2.5
2012
88
Evaluation year
2009
5.0
2015
Δ Macro AUC
2015
Macro AUC (%)
2.5
0.0
2012
Evaluation year
2012
7.5
2018
88 2000
94
10.0
2021
Training year
2018
Training year
2015
5.0
Δ Macro AUC
Macro AUC (%)
Training year
Training year
2018
2015
2025 2024
98
96
2018
2025 2024
−10.0
Evaluation year
96
2006
−7.5
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00
20 2024 25
20 21
20 18
20 15
20 12
20 09
20 06
20 03
20 00 2021
92
−5.0
2003
(b) DistilBERT
2025 2024
98
2009
−2.5
2006
Evaluation year
(a) BERT
2012
0.0
2009
2000
Evaluation year
94
2.5
2012
88
Evaluation year
2015
5.0
2015
Δ Macro AUC
0.0
2012
94
Macro AUC (%)
2015
88 2000
2025 2024
7.5
2018
Training year
2.5
Training year
2018
Macro AUC (%)
2006
2015
5.0
Δ Macro AUC
92 2009
Macro AUC (%)
2012
Training year
Training year
2018
94
10.0
2021
96
2018 2015
2025 2024
98
96
Training year
2025 2024
Evaluation year
(h) MiniLM-L6
Figure 26: Frozen models: Macro AUC drift matrix M (m) and deviation from the cohort mean ∆(m) = M (m) − M̄ for each model, shown on a sequential and a zero-centred diverging scale, respectively.
E.7. Forgetting and Rankings To see how quickly each model forgets, we summarize its drift matrix as a forgetting curve. The curve plots the Macro AUC against the lag ℓ = j − i, the number of slices between the training cutoff i and the evaluation slice j. At each lag we average over all training cutoffs, F (ℓ) = meani Mi, i+ℓ . The result is the Macro AUC at a fixed temporal distance, independent of which period a model was trained on. This separates the effect of temporal distance from the difficulty of any single slice. 30
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift 98 BERT BiGRU-S BiLSTM-Attn-L BiLSTM-M DeBERTa-v3 DistilBERT ELECTRA FFN-L
90
96 94 Macro AUC (%)
Macro AUC (%)
95
FFN-M FFN-S MiniLM-L6 ModernBERT MPNet RoBERTa
85
80
92
FFN Frozen Recurrent
90
TextCNN Transformer
88
TextCNN-L TextCNN-M TextCNN-S
86
TX-L TX-M
84
TX-S
82 0
5
10 15 Years since training
20
0
25
5
10 15 Years since training
(a) Per model
20
25
(b) Per model family
Figure 27: Forgetting curves: each model (left) and averaged within each family (right).
Decay
Future performance TextCNN-L
95.2
ModernBERT
BiLSTM-M
95.2
FFN-L
2.7
BiLSTM-Attn-L
95.1
TextCNN-L
2.7
TX-M
95.1
TX-M
2.8
TX-S
95.1
MiniLM-L6
2.8
ModernBERT
94.9
FFN-M
2.9
FFN-L
94.9
TX-S
2.9
MiniLM-L6
94.8
BiLSTM-Attn-L
2.9
TextCNN-M
94.8
BiLSTM-M
2.9
TX-L
94.7
TextCNN-M
3.1
BiGRU-S
94.7
BERT
3.1
FFN-M
94.6
TX-L
3.1
TextCNN-S
94.5
BiGRU-S
3.2
FFN-S
93.9
DistilBERT
3.3
BERT
93.6
FFN-S
3.4
DistilBERT
93.1
TextCNN-S
3.4
RoBERTa
92.3
RoBERTa
3.6
MPNet
92.0
MPNet
3.7
DeBERTa-v3
DeBERTa-v3
86.1
ELECTRA
20
40 60 Macro AUC (%)
6.5
ELECTRA
84.4
0
2.5
7.3
0
80
(a) By future performance
1
2
3
4 Decay (%)
(b) By decay
Figure 28: Models ranked by mean future performance and by temporal decay.
31
5
6
7
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
E.8. Result Tables Table 21: Temporal robustness on arXiv. Model ModernBERT FFN-L TextCNN-L TX-M MiniLM-L6 FFN-M TX-S BiLSTM-Attn-L BiLSTM-M TextCNN-M
Future (%)
Decay (%)
94.9 94.9 95.2 95.1 94.8 94.6 95.1 95.1 95.2 94.8
2.5 2.7 2.7 2.8 2.8 2.9 2.9 2.9 2.9 3.1
Model BERT TX-L BiGRU-S DistilBERT FFN-S TextCNN-S RoBERTa MPNet DeBERTa-v3 ELECTRA
Future (%)
Decay (%)
93.6 94.7 94.7 93.1 93.9 94.5 92.3 92.0 86.1 84.4
3.1 3.1 3.2 3.3 3.4 3.4 3.6 3.7 6.5 7.3
Table 22: arXiv: models trained up to 2000, ordered by future Table 23: arXiv: models trained up to 2008, ordered by future performance. performance. Rank
Model
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
MiniLM-L6 ModernBERT TextCNN-L TX-M BiLSTM-Attn-L TX-S TX-L TextCNN-M FFN-L TextCNN-S BiLSTM-M FFN-M BiGRU-S FFN-S BERT DistilBERT RoBERTa MPNet DeBERTa-v3 ELECTRA
Macro AUC (%)
Future (%)
Decay (%)
Rank
Model
97.9 97.5 98.1 96.2 97.8 96.6 96.2 97.7 97.2 97.4 97.3 96.6 93.3 95.9 95.8 94.6 91.5 88.9 90.7 88.9
93.9 93.9 93.8 93.6 93.5 93.5 93.4 93.4 93.2 93.0 92.8 92.3 91.9 91.3 90.9 90.2 88.0 87.3 81.2 77.3
4.0 3.7 4.3 2.7 4.3 3.2 2.8 4.3 4.0 4.4 4.5 4.3 1.4 4.6 4.9 4.4 3.5 1.6 9.5 11.7
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
TX-M TextCNN-L BiLSTM-Attn-L BiLSTM-M TX-S FFN-L ModernBERT MiniLM-L6 TextCNN-M FFN-M TX-L BiGRU-S BERT FFN-S TextCNN-S DistilBERT RoBERTa MPNet DeBERTa-v3 ELECTRA
Macro AUC (%)
Future (%)
Decay (%)
98.6 98.7 98.5 98.8 98.4 98.0 98.0 97.9 98.6 97.9 98.2 98.6 97.4 97.6 98.5 97.1 96.8 96.7 93.4 92.5
96.0 95.7 95.6 95.6 95.3 95.2 95.0 94.9 94.9 94.9 94.8 94.5 93.8 93.7 93.6 93.3 92.5 92.5 86.1 84.6
2.6 3.0 2.9 3.2 3.1 2.8 3.0 3.0 3.7 3.0 3.4 4.2 3.6 4.0 4.9 3.9 4.3 4.3 7.3 7.9
Table 24: arXiv: models trained up to 2016, ordered by future Table 25: arXiv: models trained up to 2024, ordered by future performance. performance. Rank
Model
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
BiLSTM-M BiGRU-S TX-S TX-M TX-L BiLSTM-Attn-L TextCNN-L TextCNN-S TextCNN-M FFN-L FFN-M MiniLM-L6 FFN-S ModernBERT BERT DistilBERT RoBERTa MPNet DeBERTa-v3 ELECTRA
Macro AUC (%)
Future (%)
Decay (%)
Rank
Model
98.8 98.8 98.7 98.7 98.7 98.7 98.6 98.5 98.5 98.3 98.2 98.4 98.1 98.2 97.6 97.3 97.0 97.0 93.6 92.6
96.6 96.6 96.5 96.5 96.5 96.4 96.3 96.2 96.2 96.2 96.1 96.0 96.0 95.8 95.1 94.8 94.6 94.4 90.0 88.3
2.2 2.2 2.2 2.2 2.2 2.3 2.3 2.3 2.4 2.2 2.2 2.4 2.2 2.4 2.5 2.5 2.4 2.5 3.7 4.3
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
TX-M BiLSTM-M TX-L BiLSTM-Attn-L BiGRU-S TX-S TextCNN-S FFN-L FFN-M TextCNN-M FFN-S MiniLM-L6 TextCNN-L ModernBERT BERT DistilBERT RoBERTa MPNet DeBERTa-v3 ELECTRA
32
Macro AUC (%)
Future (%)
Decay (%)
96.5 96.5 96.4 96.4 96.4 96.4 96.2 96.0 96.0 96.0 95.9 95.9 95.8 95.6 95.1 94.9 94.9 94.9 92.1 91.5
96.3 96.3 96.3 96.3 96.3 96.2 96.0 95.8 95.8 95.8 95.7 95.6 95.5 95.2 94.7 94.5 94.5 94.4 91.2 90.4
0.1 0.2 0.2 0.2 0.2 0.2 0.2 0.2 0.2 0.3 0.2 0.3 0.3 0.4 0.4 0.4 0.4 0.4 0.8 1.1
Drift Happens: Neural Architecture Robustness to Temporal Distribution Shift
Table 26: arXiv: future performance and decay by model family.
Family FFN Frozen Recurrent TextCNN Transformer
2000 Future Decay 92.2 87.8 92.7 93.4 93.5
4.3 5.4 3.4 4.3 2.9
2008 Future Decay 94.6 91.6 95.2 94.7 95.4
3.3 4.6 3.4 3.9 3.0
33
2016 Future Decay 96.1 93.6 96.5 96.2 96.5
2.2 2.8 2.2 2.3 2.2
2024 Future Decay 95.8 93.8 96.3 95.8 96.3
0.2 0.5 0.2 0.3 0.2