ConceptioArchivearXiv CS
arXiv CSopen access

In-Context Time Series Classification with Random Convolutional Features

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

In-Context Time Series Classification with Random Convolutional Features Joscha Cüppers1 , Jilles Vreeken1 1

CISPA Helmholtz Center for Information Security [email protected], [email protected]

arXiv:2607.19234v1 [cs.LG] 21 Jul 2026

Abstract Time series classification is central to domains like medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-channel interactions. Random convolutional transforms efficiently map these sequences to fixed-dimensional tabular features but are traditionally paired with simple linear classifiers. We investigate whether a pretrained tabular foundation model can more effectively harness these rich representations. We propose MASHT, a pipeline that marries MultiRocket and Hydra features with the power of in-context tabular foundation models. By leveraging a pretrained tabular foundation model, our approach completely bypasses task-specific model training, requiring only feature extraction and direct inference. Extensive experiments demonstrate that MASHT matches state-of-the-art time series classification baselines on univariate tasks, achieving a lower average rank than HIVECOTE 2.0. On multivariate datasets, MASHT remains highly competitive with the strongest reference methods. GitHub: https://github.com/joschac/masht

1

Introduction

Time series classification (TSC) appears in domains ranging from healthcare and human activity recognition to manufacturing and infrastructure monitoring. Its central difficulty is representational: class information may be expressed as local shapes, frequencies, phase-independent motifs, level changes, or interactions between channels (Bagnall et al. 2017; Ismail Fawaz et al. 2019). Highly accurate ensemble methods such as HIVE-COTE 2.0 cover many such representations, but they often achieve this breadth at substantial computational cost (Middlehurst et al. 2021). Random convolutional transforms offer a complementary design point. ROCKET (Dempster, François, and Webb 2020) showed that fixed convolutional kernels paired with a simple classifier can classify accurately and efficiently. MiniRocket and MultiRocket made this family faster and more accurate (Dempster, Schmidt, and Webb 2021; Tan et al. 2022), while Hydra added competing convolutional count features with a dictionary-like interpretation (Dempster, Schmidt, and Webb 2023). Together, MultiRocket and Hydra provide a provide broad temporal features.

The downstream classifier for these transformed tables is typically simple, like ridge or linear classifiers. They offer speed and robustness, but they fail to capture nonlinear interactions between the transformed features. In contrast, TabPFN-3 is a tabular foundation model pretrained to perform in-context prediction on entirely new supervised tables without any downstream training (Grinsztajn et al. 2026). This suggests a modular route to TSC: use specialized temporal transforms to build a tabular features, then use a pretrained tabular model for dataset-conditioned classification. This paper introduces MASHT (MultiRocket And Stacked Hydra Transformed), a modular pipeline that connects efficient temporal representations with pretrained, in-context tabular classification to improve time series classification (TSC). By concatenating random convolutional representations from MultiRocket and Hydra, the pipeline feeds a comprehensive feature table into TabPFN-3 for classification. We evaluate MASHT against established univariate and multivariate TSC baselines across 112 univariate and 71 multivariate datasets using aggregate benchmarks, mean ranks, and Holm-corrected pairwise Wilcoxon significance tests. To complement these aggregate results, we examine datasetlevel performance and runtime, highlighting when the representation is effective and where its limitations emerge.

2

Related Work

Our work connects two research directions: efficient representations for time series classification and in-context prediction with tabular foundation models. We first situate the approach within the broader TSC literature and its benchmarking practices, then review the random convolutional and dictionary-like transforms used to construct fixed tabular features, and finally discuss tabular foundation models as downstream classifiers for these representations.

Time Series Classification Time series classification (TSC) methods are commonly grouped into distance-, dictionary-, interval-, shapelet-, kernel-, feature-, and deep-learning-based approaches (Bagnall et al. 2017; Ismail Fawaz et al. 2019). HIVE-COTE 2.0 combines multiple representations in a large hierarchical ensemble and remains a prominent high-accuracy reference point (Middlehurst et al. 2021). The broader benchmarking

9

8

7

6

5

4

3

PF 7.0536 FreshPRINCE 6.5357 QUANT 5.4196 WEASEL-2 5.3705

2

1

3.2723 MASHT 3.4062 HC2 3.8839 MR-Hydra 4.9464 H-InceptionTime 5.1116 RDST

Figure 1: Critical-difference diagram for accuracy on UTF-112. Methods connected by a horizontal bar are not significantly different under the Wilcoxon–Holm comparison.

literature provides the context for both univariate and multivariate comparisons (Dau et al. 2019; Ruiz et al. 2021; Middlehurst et al. 2026).

Efficient Time Series Feature Transforms ROCKET transforms each series with random convolutional kernels and summarizes the responses using pooling operators (Dempster, François, and Webb 2020). MiniRocket restricts the kernel construction to obtain an almost deterministic and much faster transform (Dempster, Schmidt, and Webb 2021). MultiRocket expands this design with multiple pooling operators and features from both the raw series and its first-order difference (Tan et al. 2022). Hydra organizes convolutional kernels into competing groups and records which kernels produce the strongest responses (Dempster, Schmidt, and Webb 2023). Its features therefore capture dictionary-like counts while retaining efficient convolutional computation. Prior work reports that concatenating Hydra with ROCKET-family features improves accuracy, motivating the combined representation used here. QUANT and WEASEL 2.0 provide complementary interval and dictionary perspectives on scalable TSC (Dempster, Schmidt, and Webb 2024; Schäfer and Leser 2023).

Tabular Foundation Models TabPFN treats supervised prediction as inference conditioned on a labeled training table. Rather than fitting a model from scratch for every dataset, it uses a transformer pretrained on synthetic tasks to implement a learned prediction algorithm (Hollmann et al. 2025). TabPFN-3 extends this line to greater scale and faster inference (Grinsztajn et al. 2026). Our work sits between these literatures. Prior TSC work studies strong temporal representations, and prior tabular foundation-model work studies generic tables. This paper tests whether a pretrained tabular classifier can improve fixed random-convolutional TSC representations after temporal structure has been encoded by MultiRocket and Hydra.

3

Method

MASHT follows a two-stage design. It first maps each input series to a fixed-dimensional feature table by concatenat-

ing complementary MultiRocket (Tan et al. 2022) and Hydra (Dempster, Schmidt, and Webb 2023) representations. It then supplies the transformed training examples and query series to the pretrained TabPFN-3 (Grinsztajn et al. 2026) classifier for in-context prediction, without learning a task-specific temporal model end to end. This separation lets the random convolutional transforms encode temporal structure while the tabular foundation model captures relationships among the resulting features. We next formalize the prediction problem before describing the pipeline and its feature-budgeted implementation.

Problem Setting Let Dtrain = {(xi , yi )}ni=1 be a labeled TSC dataset. A univariate series is xi ∈ RL and a multivariate series is xi ∈ RC×L , where C is the number of channels and L is the series length. The goal is to estimate class probabilities p(y | x∗ , Dtrain ) for an unseen series x∗ .

Pipeline The proposed pipeline computes a MultiRocket representation ϕMR (x), a Hydra representation ϕH (x), and then classifies their concatenation: z(x) = [ϕMR (x); ϕH (x)], p̂(y | x∗ ) = fTabPFN-3 (z(x∗ ); Dz ),

(1)

where Dz = {(z(xi ), yi )}ni=1 is the transformed training set.

Feature Extraction We use a dataset-dependent feature budget to balance representation richness against transformation, memory, and inference costs. Because both transforms generate features in discrete blocks, the realized dimensionality may be slightly below this target. MultiRocket applies fixed convolutional kernels over both the raw input and its first-order difference. Kernel responses are summarized by multiple pooling operators, including proportions, positive magnitudes, positive-response locations, and stretch lengths (Tan et al. 2022). The requested MultiRocket budget is split across raw and differenced inputs, with four pooling features per kernel.

Hydra samples groups of competing convolutional kernels at multiple dilations and counts the kernels attaining extreme responses, by default, maximum and minimum counts (Dempster, Schmidt, and Webb 2023). Each group contains k kernels evaluated at the same dilation. At every temporal position, the kernels within a group compete, and Hydra aggregates the winning max and min responses into separate features for each kernel. Each group can hence be viewed as a small dictionary of random patterns, with multiple groups providing independent dictionaries at every dilation. Hydra adapts the kernel dilation to the length of the series L, which results in more features for longer series. To keep the number of features below the Hydra budget B, we therefore vary the number of groups g. Hydra uses powers-of-two dilations up to the maximum length per series; the number (L−1) of dilations per group is hence d = log29−1 , the term 9 − 1 comes from Hydra’s fixed kernel length of 9. For a requested Hydra budget B, we therefore set the number of groups to,   B g= . 2kd The factor of 1/2 accounts for the two Hydra feature blocks produced per group and dilation, corresponding to maximum- and minimum-response counts. We fix k = 8 kernels per group, as per default in Hydra.This adjustment keeps the Hydra representation within the feature budget.

4

Evaluation

We evaluate MASHT on established univariate and multivariate time-series classification benchmarks. The reference results are taken from the corresponding studies, and obtain all MASHT results under the same dataset protocols.

Experiment Setup We adapt the feature budget to the combined number of training and test instances, using 10,000 features for fewer than 1,000 instances, 2,000 for fewer than 100,000, and 200 otherwise. The budget is divided equally between Hydra (Dempster, Schmidt, and Webb 2023) and MultiRocket (Tan et al. 2022) before their features are concatenated. We provide further implementation details in Appendix A. Evaluation We primarily evaluate using accuracy. On UTF-112, we also report balanced accuracy and AUROC. Balanced accuracy is the unweighted mean of the recall obtained for each class, so every class contributes equally regardless of its frequency. For binary datasets, we compute AUROC with the minority class in the training split as the positive class. For multiclass datasets, we compute a one-vsrest AUROC for each class and average the resulting scores using the class frequencies in the training split as weights. For the univariate and multivariate benchmark, we rank methods on every dataset and report their mean ranks. Pairwise comparisons use the two-sided Wilcoxon signedrank test with Holm correction (Holm 1979) and report wins/draws/losses together with the mean accuracy difference. Higher values are better for all reported metrics, while lower mean ranks are better.

Table 1: Mean performance on the 112 univariate UTF-112 datasets. Baseline results are from Middlehurst, Schäfer, and Bagnall (2024); MASHT results are on the same 30 resamples. Rank is the mean per-dataset accuracy rank. Method

Acc.

Bal. Acc.

AUROC

Rank

MASHT HC2 MR-Hydra H-InceptionTime RDST WEASEL-2 QUANT FreshPRINCE PF

0.892 0.891 0.884 0.876 0.876 0.874 0.867 0.855 0.837

0.872 0.871 0.866 0.861 0.856 0.853 0.845 0.834 0.819

0.970 0.968 0.913 0.959 0.907 0.905 0.962 0.958 0.942

3.27 3.41 3.88 4.95 5.11 5.37 5.42 6.54 7.05

Univariate Evaluation We follow the UTF-112 evaluation of Middlehurst, Schäfer, and Bagnall (2024), comprising 112 univariate datasets and 30 resamples per dataset. We evaluate MASHT on the same resamples1 and compare it with the results reported by Middlehurst, Schäfer, and Bagnall (2024). The comparison includes HIVE-COTE 2.0 (HC2) (Middlehurst et al. 2021), MR-Hydra (Tan et al. 2022; Dempster, Schmidt, and Webb 2023), QUANT (Dempster, Schmidt, and Webb 2024), RDST (Guillaume, Vrain, and Elloumi 2022), WEASEL-2 (Schäfer and Leser 2023), FreshPRINCE (Middlehurst and Bagnall 2022), Proximity Forest (PF) (Lucas et al. 2019), and HInceptionTime (Ismail Fawaz et al. 2020). In Table 1 we show the mean aggregate results over 30 resamples. We observe that MASHT achieves the highest mean accuracy, balanced accuracy, and AUROC, as well as the best mean accuracy rank. The margin over HC2 is small: MASHT improves mean accuracy from 0.891 to 0.892. In Figure 1, we show the critical-difference plot of the perdataset accuracy ranks. We observe no statistically significant differences between MASHT and HC2, however MASHT is statistically significant better than MR-Hydra. In Table 2, we provide a more detailed pairwise comparison of MASHT against each baseline, reporting wins, draws, losses, mean accuracy differences, and Holm-adjusted p-values. Against HC2, MASHT wins on 58 datasets, draws on 5, and loses on 49; the mean accuracy difference is 0.001, and the difference is not significant after Holm correction. The comparison with MR-Hydra is particularly relevant because it uses the same MultiRocket–Hydra representation family without the TabPFN-3 classifier. Here, MASHT wins on 65 datasets, draws on 7, and loses on 40, with a mean difference of 0.008; this difference remains significant after Holm correction. In Figure 2, we show the per-dataset accuracy of MASHT 1

The authors provide the resampled dataset at https://tsml-eval.readthedocs.io/en/stable/publications/2023/tsc_ bakeoff/tsc_bakeoff_2023.html — Direkt link to data: https://drive. google.com/file/d/1V36LSZLAK6FIYRfPx6mmE5euzogcXS83/ view?usp=sharing

Table 2: Pairwise accuracy comparison of MASHT with all compared UTF-112 baselines. W/D/L denotes the numbers of datasets on which MASHT wins, draws, or loses. Differences are computed as MASHT minus the baseline. Baseline

W/D/L

Mean ∆

p

Holm p

HC2 MR-Hydra H-InceptionTime RDST WEASEL-2 QUANT FreshPRINCE PF

58/5/49 65/7/40 69/4/39 75/5/32 79/3/30 86/5/21 97/3/12 94/5/13

0.001 0.008 0.016 0.016 0.018 0.025 0.037 0.055

0.42 0.01 0.00 0.00 0.00 0.00 0.00 0.00

0.42 0.03 0.00 0.00 0.00 0.00 0.00 0.00

ICTS better

MASHT accuracy

1.0

HC2 (58/5/49)

Equal

Other method better

MR-Hydra (65/7/40)

0.8

Method

Acc.

Rank

HC2 MR-Hydra MASHT RDST FreshPRINCE Arsenal CIF-500 ROCKET QUANT DrCIF-500 RIST litetime-mv STC h-inceptiontime Catch22 TDE 1NN-DTW

0.805 0.797 0.795 0.797 0.792 0.792 0.796 0.793 0.786 0.769 0.778 0.750 0.784 0.731 0.756 0.759 0.683

6.51 7.54 7.65 7.66 7.99 8.01 8.20 8.23 8.88 8.94 8.95 9.70 10.13 10.43 10.60 10.81 12.76

0.6 0.4 0.50

0.75

HC2 accuracy

1.00

0.50

0.75

MR-Hydra accuracy

1.00

Figure 2: Per-dataset UTF-112 accuracy comparisons with MASHT. Panel titles report wins/draws/losses for MASHT.

against HC2 and MR-Hydra, and see how these aggregate differences arise. Against HC2, the dataset-level results cluster closely around the diagonal, indicating very similar performance without systematic large differences. The comparison with MR-Hydra is also close for the majority of datasets, but MASHT performs clearly better on a small number of datasets. Thus, the improvement over MR-Hydra is concentrated in several pronounced gains rather than a uniform advantage across the benchmark.

Multivariate Evaluation For multivariate classification, we follow the Multiverse benchmark of Middlehurst et al. (2026),2 which comprises 71 datasets. We evaluate MASHT on these and compare it with the results reported by Middlehurst et al. (2026). The benchmark covers a broad selection of distance-based, interval, shapelet, dictionary, feature-based, convolutional, hybrid, and deep-learning classifiers, including HC2 (Middlehurst et al. 2021), MR-Hydra (Tan et al. 2022; Dempster, Schmidt, and Webb 2023), QUANT (Dempster, Schmidt, and Webb 2024), ROCKET (Dempster, François, and Webb 2

Table 3: Mean accuracy on the 71 multivariate Multiverse datasets. Baseline results are from Middlehurst et al. (2026). Rank is the mean per-dataset accuracy rank.

We compare on all datasets where the results for all methods are available https://github.com/aeon-toolkit/multiverse/blob/ main/results/multiverse/accuracy_mean.csv

Table 4: Pairwise accuracy comparison of MASHT with selected Multiverse baselines. W/D/L denotes the numbers of datasets on which MASHT wins, draws, or loses. Differences are computed as MASHT minus the baseline. Baseline HC2 MR-Hydra RDST QUANT ROCKET Arsenal FreshPRINCE

W/D/L

Mean ∆

p

Holm p

23/14/34 23/16/32 33/7/31 35/14/22 33/9/29 31/11/29 35/12/24

-0.010 -0.002 -0.002 0.010 0.002 0.003 0.004

0.15 0.45 0.84 0.39 0.53 0.65 0.61

1.00 1.00 1.00 1.00 1.00 1.00 1.00

2020), Arsenal (Middlehurst et al. 2021), FreshPRINCE (Middlehurst and Bagnall 2022), CIF (Middlehurst, Large, and Bagnall 2020), DrCIF (Middlehurst, Large, and Bagnall 2020), RDST (Guillaume, Vrain, and Elloumi 2022), RIST (Middlehurst and Bagnall 2023), STC (Bostrom and Bagnall 2017), TDE (Middlehurst et al. 2020), Catch22 (Lubba et al. 2019), InceptionTime (Ismail Fawaz et al. 2020), LITETime (Ismail-Fawaz et al. 2025), and 1NN-DTW (Shokoohi-Yekta et al. 2017). In Table 3, we report the mean accuracy and rank of each method. In Figure 3, we show the corresponding criticaldifference diagram. We observe that HC2 obtains the best mean rank and mean accuracy, followed by MR-Hydra, with MASHT close behind. MASHT reaches a mean accuracy of 0.795, compared with 0.805 for HC2 and 0.797 for MRHydra. However, the leading methods are not significantly different. Overall, the multivariate results are more conservative than the univariate results. In Table 4, we report pairwise accuracy comparisons of

17

16

15

14

13

12

11

10

9

8

7

6

5

4

3

2

1NN-DTW 12.7606 TDE 10.8099 Catch22 10.5986 h-inceptiontime 10.4296 STC 10.1338 litetime-mv 9.7042 RIST 8.9507 DrCIF-500 8.9437

1

6.5070 7.5352 7.6479 7.6620 7.9930 8.0141 8.1972 8.2324 8.8803

HC2 MR-Hydra MASHT RDST FreshPRINCE Arsenal CIF-500 ROCKET QUANT

Figure 3: Critical-difference diagram for accuracy on the Multiverse benchmark. ICTS better 1.00

MASHT accuracy

Equal

HC2 (23/14/34)

5

Other method better

MR-Hydra (23/16/32)

The results reveal a benchmark-dependent picture that is not captured by aggregate accuracy alone. We therefore discuss what the contrast between the univariate and multivariate evaluations suggests about the representation, place the observed performance in the context of runtime, and outline the main limitations of the current study.

0.75 0.50 0.25 0.00

Discussion

Benchmark-Dependent Performance 0.0

0.5

HC2 accuracy

1.0 0.0

0.5

MR-Hydra accuracy

1.0

Figure 4: Per-dataset Multiverse accuracy comparisons with MASHT. Panel titles report wins/draws/losses for MASHT.

MASHT with selected Multiverse baselines. We observe that, against HC2, MASHT wins on 23 datasets, draws on 14, and loses on 34, with a mean accuracy difference of −0.010. Against MR-Hydra, it wins on 23, draws on 16, and loses on 32, with a mean difference of −0.002. None of the reported pairwise comparisons is significant. These small differences show that MASHT is competitive with the strongest methods on the Multiverse benchmark, but does not establish a new state of the art there. In Figure 4, we show the per-dataset accuracy of MASHT against HC2 and MR-Hydra. We observe larger method-tomethod performance differences than in the univariate comparisons in Figure 2, with the points more widely dispersed around the diagonal. MASHT performs particularly well on several binary or low-class, and short high-channel motion datasets such as the IRDS and UIPRMD variants. HC2 more often leads on biosignal, gesture, and motion datasets with richer class structure or longer temporal context, including Handwriting, IEEEPPG, EthanolConcentration, and PhotoStimulation. Although these patterns are not causal evidence, they suggest that MASHT is strongest when its random convolutional features provide a compact tabular representation that TabPFN-3 can exploit, whereas HC2 benefits from combining a broader range of temporal representations.

The contrast between the two benchmarks is the central empirical finding. MASHT performs in the leading accuracy tier on UTF-112 and improves on MR-Hydra, whereas on Multiverse it is competitive but does not lead. This suggests that representing a time series as a compact feature table is particularly effective in the univariate setting. The comparison with MR-Hydra further indicates that TabPFN-3 is a promising classifier for the MultiRocket–Hydra features. The multivariate results also indicate where the current representation may be less effective. HIVE-COTE 2.0 combines interval, dictionary, shapelet, and convolutional representations, which may capture complementary dependencies between channels and across time. By comparison, MASHT asks a tabular classifier to recover predictive structure from a fixed random-convolutional representation. This explanation is consistent with the observed benchmark contrast, but remains a hypothesis for future study.

Runtime Runtime provides an important additional perspective on the accuracy results. Our median end-to-end runtime for MASHT on UTF-112 is approximately 59.7 seconds per dataset (Table 5). For comparison, Middlehurst, Schäfer, and Bagnall (2024) report a median HC2 training time of 15.28 minutes across 142 univariate datasets, with a total training time of 263.89 hours and individual times ranging from 38.94 seconds to 65.66 hours. Their experiments used an Intel Xeon Gold 5220R CPU, whereas all MASHT experiments were run using an NVIDIA A100 GPU with 40 GB of memory. The reported timing boundaries and dataset collections also differ. The values therefore suggest that MASHT may offer a favorable accuracy–runtime trade-off, but they do not constitute a controlled speed comparison; establishing one requires both methods to be run on the same datasets, hardware, and timing protocol.

Table 5: Runtime summary for MASHT in seconds per dataset. Full runtime comprises feature extraction and classification. Runtime component UTF-112 Total runtime Hydra transform MultiROCKET transform Classification Multiverse Total runtime Hydra transform MultiROCKET transform Classification

Mean (s)

Median (s)

Maximum (s)

43.954 0.939 0.085 42.928

59.743 0.456 0.059 59.528

83.457 7.884 0.507 83.124

47.312 1.177 0.279 45.854

57.813 0.375 0.100 57.144

79.962 14.173 4.848 70.312

Limitations The method relies on a pretrained foundation model with nontrivial inference and memory requirements. Finally, the conclusions are limited to the datasets and evaluation protocols of UTF-112 and Multiverse; broader claims require additional benchmarks and application-specific evaluations.

6

Conclusion

We presented MASHT, a modular time-series classification approach that combines MultiRocket and Hydra features with TabPFN-3. On the UTF-112 benchmark, MASHT achieves the highest mean accuracy and significantly outperforms MR-Hydra, while remaining statistically tied with HIVECOTE 2.0. On the Multiverse benchmark, it is competitive with the strongest published methods but does not lead them. These results show that tabular foundation models are a promising classifier for random convolutional time-series features, particularly in the univariate setting. Standardized efficiency comparisons are the most important next steps.

References Bagnall, A.; Lines, J.; Bostrom, A.; Large, J.; and Keogh, E. 2017. The Great Time Series Classification Bake Off: A Review and Experimental Evaluation of Recent Algorithmic Advances. Data Mining and Knowledge Discovery, 31(3): 606–660. Bostrom, A.; and Bagnall, A. 2017. Binary Shapelet Transform for Multiclass Time Series Classification. Transactions on Large-Scale Data- and Knowledge-Centered Systems, 32: 24–46. Dau, H. A.; Bagnall, A.; Kamgar, K.; Yeh, C.-C. M.; Zhu, Y.; Gharghabi, S.; Ratanamahatana, C. A.; and Keogh, E. 2019. The UCR Time Series Archive. IEEE/CAA Journal of Automatica Sinica, 6(6): 1293–1305. Dempster, A.; François, P.; and Webb, G. I. 2020. ROCKET: exceptionally fast and accurate time series classification using random convolutional kernels. Data Mining and Knowledge Discovery, 34(5): 1454–1495. Dempster, A.; Schmidt, D. F.; and Webb, G. I. 2021. MiniRocket: A Very Fast (Almost) Deterministic Transform for Time Series Classification. In Proceedings of the 27th

ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 248–257. Dempster, A.; Schmidt, D. F.; and Webb, G. I. 2023. Hydra: Competing Convolutional Kernels for Fast and Accurate Time Series Classification. Data Mining and Knowledge Discovery, 37: 1779–1805. Dempster, A.; Schmidt, D. F.; and Webb, G. I. 2024. quant: a minimalist interval method for time series classification. Data Mining and Knowledge Discovery, 38(4): 2377–2402. Grinsztajn, L.; Flöge, K.; Key, O.; Birkel, F.; Jund, P.; Roof, B.; Manium, M.; Hoo, S. B.; Bühler, M.; Garg, A.; et al. 2026. TabPFN-3: Technical Report. arXiv preprint arXiv:2605.13986. Guillaume, A.; Vrain, C.; and Elloumi, W. 2022. Random dilated shapelet transform: A new approach for time series shapelets. In International Conference on Pattern Recognition and Artificial Intelligence, 653–664. Springer. Hollmann, N.; Muller, S.; Purucker, L.; Krishnakumar, A.; Korfer, M.; Hoo, S. B.; Schirrmeister, R. T.; and Hutter, F. 2025. Accurate Predictions on Small Data with a Tabular Foundation Model. Nature, 637: 319–326. Holm, S. 1979. A simple sequentially rejective multiple test procedure. Scandinavian journal of statistics, 65–70. Ismail-Fawaz, A.; Devanne, M.; Berretti, S.; Weber, J.; and Forestier, G. 2025. Look into the LITE in Deep Learning for Time Series Classification. International Journal of Data Science and Analytics, 20: 4029–4049. Ismail Fawaz, H.; Forestier, G.; Weber, J.; Idoumghar, L.; and Muller, P.-A. 2019. Deep Learning for Time Series Classification: A Review. Data Mining and Knowledge Discovery, 33(4): 917–963. Ismail Fawaz, H.; Lucas, B.; Forestier, G.; Pelletier, C.; Schmidt, D. F.; Weber, J.; Webb, G. I.; Idoumghar, L.; Muller, P.-A.; and Petitjean, F. 2020. InceptionTime: Finding AlexNet for Time Series Classification. Data Mining and Knowledge Discovery, 34(6): 1936–1962. Lubba, C. H.; Sethi, S. S.; Knaute, P.; Schultz, S. R.; Fulcher, B. D.; and Jones, N. S. 2019. catch22: CAnonical Timeseries CHaracteristics selected through highly comparative time-series analysis. bioRxiv. Lucas, B.; Shifaz, A.; Pelletier, C.; O’Neill, L.; Zaidi, N.; Goethals, B.; Petitjean, F.; and Webb, G. I. 2019. Proximity Forest: An Effective and Scalable Distance-Based Classifier for Time Series. Data Mining and Knowledge Discovery, 33(3): 607–635. Middlehurst, M.; and Bagnall, A. 2022. The FreshPRINCE: A Simple Transformation Based Pipeline Time Series Classifier. In International Conference on Pattern Recognition and Artificial Intelligence, 150–161. Middlehurst, M.; and Bagnall, A. 2023. Extracting features from random subseries: A hybrid pipeline for time series classification and extrinsic regression. In International Workshop on Advanced Analytics and Learning on Temporal Data, 113–126. Springer.

Middlehurst, M.; Large, J.; and Bagnall, A. 2020. The Canonical Interval Forest (CIF) Classifier for Time Series Classification. In IEEE International Conference on Big Data, 188–195. Middlehurst, M.; Large, J.; Cawley, G.; and Bagnall, A. 2020. The temporal dictionary ensemble (TDE) classifier for time series classification. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 660–676. Springer. Middlehurst, M.; Large, J.; Flynn, M.; Lines, J.; Bostrom, A.; and Bagnall, A. 2021. HIVE-COTE 2.0: a new meta ensemble for time series classification. Machine Learning, 110(11): 3211–3243. Middlehurst, M.; Rushbrooke, A.; Ismail-Fawaz, A.; Devanne, M.; Forestier, G.; Dempster, A.; Webb, G. I.; Holder, C.; and Bagnall, A. 2026. The Multiverse of Time Series Machine Learning: An Archive for Multivariate Time Series Classification. arXiv preprint arXiv:2603.20352. Middlehurst, M.; Schäfer, P.; and Bagnall, A. 2024. Bake off redux: a review and experimental evaluation of recent time series classification algorithms. Data Mining and Knowledge Discovery, 38(4): 1958–2031. Ruiz, A. P.; Flynn, M.; Large, J.; Middlehurst, M.; and Bagnall, A. 2021. The Great Multivariate Time Series Classification Bake Off: A Review and Experimental Evaluation of Recent Algorithmic Advances. Data Mining and Knowledge Discovery, 35(2): 401–449. Schäfer, P.; and Leser, U. 2023. WEASEL 2.0: A Random Dilated Dictionary Transform for Fast, Accurate and Memory Constrained Time Series Classification. Machine Learning, 112(12): 4763–4788. Shokoohi-Yekta, M.; Hu, B.; Jin, H.; Wang, J.; and Keogh, E. 2017. Generalizing DTW to the multi-dimensional case requires an adaptive approach. Data mining and knowledge discovery, 31(1): 1–31. Tan, C. W.; Dempster, A.; Bergmeir, C.; and Webb, G. I. 2022. MultiRocket: Multiple Pooling Operators and Transformations for Fast and Effective Time Series Classification. Data Mining and Knowledge Discovery, 36: 1623–1646.

A

Implementation Details

The benchmark code is written in Python and depends on aeon==1.4.0, tabpfn, scikit-learn, and PyTorch. The UTF-112 runs process the numbered train/test splits available for each dataset and average metrics across splits. All of our experiments were run on an NVIDIA A100 GPU with 40 GB of memory. The classifier is TabPFN-3 via the tabpfn package. The reported runs use TabPFNClassifier with 8 estimators, automatic estimator scaling, fit_mode=low_memory, CUDA execution, automatic inference precision, 8 preprocessing jobs, no tuning configuration, hidden progress bars, and ignore_pretraining_limits=True.

Record · ID 386866 · SHA-256 38d31ea6ef590b6a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.