Adaptive Anisotropic Attention for Axis-Structured Signals
arXiv:2609.08788v1 [cs.LG] 8 Sep 2026
Mahir Jain Mannas AI [email protected]
Parshva Runwal Mannas AI [email protected]
Aditya Ray Mishra Mannas AI [email protected]
Siddharth Panwar Mannas AI [email protected]
Arvasu Kulkarni Mannas AI [email protected]
Sandeep Singh Mannas AI [email protected]
Abstract Dense self-attention treats all token pairs as equally plausible before learning—an interaction-isotropic prior that can be mismatched to structured signals. For structured, low signal-to-noise ratio (SNR) signals such as EEG, dependencies are organized along the electrode and time axes, and this uniform prior exposes each token to many irrelevant interactions. We introduce Adaptive Anisotropic Attention (AAA), which splits attention into two paths: a temporal path, where each token attends to the tokens of its own electrode across time, and a spatial path, where it attends to the tokens of the other electrodes at the same time step. A small gate predicts, for every token, a convex combination and the token’s update is the weighted sum of the two path outputs. On six EEG downstream tasks, the resulting model, AXON (AXis-factorized Operator Network), improves mean balanced accuracy over a dense baseline under both linear probing and full fine-tuning. We show that both paths (temporal and spatial) are necessary and that the weighted sum beats a hard choice of one path; most of the benefit comes from the gate learning a different temporal/spatial balance at each layer of the network. Controlled audio spectrogram experiments show that axis factorization transfers beyond EEG. These results suggest that aligning attention with the natural axes of structured signals provides a useful inductive bias.
1
Introduction
Transformers [Vaswani et al., 2017] process tokens: small input segments represented as vectors. In EEG and audio, tokens form two-dimensional grids. An EEG token contains one second of signal from one electrode; an audio token covers one frequency band over one spectrogram frame. Dense self-attention connects all tokens directly, allowing long-range interactions without distinguishing the electrode–time or frequency–time axes. For structured spatiotemporal signals such as EEG, which lie on electrodes × time, this interactionisotropic prior may be mismatched. In masked autoencoding (MAE) [He et al., 2022], the model learns to reconstruct masked patches, not to distinguish downstream classes. Our hypothesis is that dense global attention encourages reconstruction shortcuts: for example, estimating a masked patch from a broad average across electrodes and time. This dilutes localized or axis-specific signals that are critical for downstream tasks such as motor imagery classification. We introduce Adaptive Anisotropic Attention (AAA), which replaces dense encoder attention with two parallel paths. The temporal path attends within each electrode across time; the spatial path attends across electrodes at the same time step. A small gate combines their outputs through a convex Preprint.
combination. This soft mixture can vary across tokens and layers. We call the resulting model AXON (AXis-factorized Operator Network) and evaluate it on six EEG tasks under linear probing (LP) and full fine-tuning (FT), against dense baselines including a parameter-matched model. We analyse learned axis mixtures and task-dependent temporal context, and test the design on audio spectrograms as a second axis-structured modality. Axis factorization helps both modalities. Related work. Factorized attention. Axial attention [Ho et al., 2019] and divided space-time attention [Bertasius et al., 2021, Arnab et al., 2021] were built for images and video and attend along one axis at a time. The axis order is set by hand and is the same for every token and every layer, and these designs were not built for, or tested on, low-SNR signals such as EEG. AXON keeps the two axes but runs them in parallel and lets a gate set the balance per token and per layer. EEG foundation models. EEGPT, LaBraM, CBraMod, BIOT, CSBrain and REVE [Wang et al., 2024, Jiang et al., 2024, Wang et al., 2025, Yang et al., 2023, Zhou et al., 2025, El Ouahidi et al., 2025] pretrain large encoders on unlabelled EEG so that one model transfers to many tasks. They differ in tokenisation and pretraining objective; CBraMod, attends along time and along channels in parallel and combines the two with a fixed split of attention heads. None of them tests whether the dense all-to-all path should be removed, or lets the time/channel balance be learned per token and per layer. We evaluate all six under one protocol (Tables 3 and 4).
2
Method: Adaptive Anisotropic Attention
EEG tasks require different temporal and spatial contexts (Appendix A). The encoder is a stack of 22 identical blocks, which we call layers; each layer has its own attention paths and its own gate weights, so “per layer” below means a separate value in each of the 22 blocks. 2.1
Input mannas.ai
We process non-overlapping 10-second windows during pretraining and task-specific windows of 4–30 seconds downstream (Table 8). Each window contains C available electrodes and T patches per electrode. A token i = (ci , ti ) is a channel–time patch on the grid Ω = C × T , where C and T index electrodes and temporal patches, respectively. The full grid contains CT tokens (231 for the pretraining grid of 21 electrodes and 11 patches). Each electrode has a known 3D head coordinate pc ∈ R3 from the standard 10–20 montage [Jasper, 1958], and xi ∈ Rd denotes the token embedding after patch projection and positional encoding. Dense attention uses shared Q, K, V projections across all tokens: ⊤ X qi kj Attn(x)i = aij V xj , aij ∝ exp √ , dh j∈Ω giving (CT )2 token pairs per layer. Dense-L widens this baseline to match AXON’s parameter count(by increasing dimensions) (Appendix C.2); And Divided-ST applies temporal then spatial attention within each block [Bertasius et al., 2021] (Appendix F.2). 2.2
Factorized attention
The core idea is to replace the dense attention operator with a weighted mixture of two axis-restricted operators, each attending within one axis of the token grid Ω = C × T : the temporal path over time within a channel, the spatial path over channels within a time step. Let T (x)i denote the temporal attention output and S(x)i the spatial attention output for token i (defined below). A single AXON block computes: yi = λ T (x)i + (1 − λ) S(x)i , λ ∈ (0, 1), where λ controls the axis mixture. In the simplest variant (AXON-Fixed), λ = σ(ℓ) is the sigmoid of a single learnable scalar ℓ per layer, shared across all tokens. The temporal and spatial paths use separate QKV and output projections (Q(T ) , K (T ) , V (T ) , O(T ) ) and (Q(S) , K (S) , V (S) , O(S) ), allowing each axis to specialise its feature space. The full cost analysis is in Appendix C.2. Although no token sees all others in one layer, on a full channel-time grid any two tokens are connected after two layers: one temporal step and one spatial step (Appendix E.1). 2
Visible EEG tokens (c, t) embeddings
S(i): same time step, all channels
AXON block (×22)
electrodes c
Temporal path T (i): same channel geometry-free, T tokens
Token gate g softmax(g(sg(xi ))/τ ) ⇒ (αi , βi )
T (i): same channel, all time steps
Spatial path S(i): same timestep content-based attn, C tokens
αi , βi
token i = (ci , ti )
Axis mixture yi = αi T (x)i + βi S(x)i
No dense global path G(x) removed (MAE shortcut)
time patches t
(a) Token grid and the two neighbourhoods.
Residual + LayerNorm + FFN
(b) AXON block.
Figure 1: (a) Channel–time grid with query token i (red), temporal neighbours T (i) (blue), and spatial neighbours S(i) (orange); dense attention uses the full grid. (b) An AXON block mixes both paths using weights from an MLP applied to sg(xi ). The encoder stacks 22 blocks without a dense global branch. Temporal path.
The temporal neighbourhood of token i is all tokens on the same channel: T (i) = {j ∈ Ω : cj = ci },
|T (i)| = T.
Each channel group of T tokens is processed as an independent sequence; channels do not interact through the temporal path. This gives each token access to the full temporal context of its electrode: short transients, rhythm-band oscillations, and longer event windows are all contained in T (i). Spatial path.
The spatial neighbourhood of token i is all tokens at the same time step: S(i) = {j ∈ Ω : tj = ti },
|S(i)| = C.
T (x)i and S(x)i are standard multi-head attention restricted to these neighbourhoods. The spatial path is not given any explicit information about electrode coordinates; electrode positions enter only through the positional encoding. (Appendix F). 2.3
Token-conditioned anisotropy gate
A fixed mixing weight shared by all tokens imposes one mixture on all of them. The token-conditioned gate lets each token choose its own. AXON-TokenGated replaces the shared layer weight with token-conditioned mixing: (αi , βi ) = softmax(g(sg(xi ))/τ ) . The two-layer MLP g : Rd → R2 has hidden width d/4, GELU activation, and a zero-initialised output layer; sg refers to stop gradient. We anneal τ from 2.0 to 1.0 over the first 1500 steps; higher values keep the weights closer to equal (Appendix C). The block output is the soft-gated mixture of the temporal and spatial paths: yi = αi T (x)i + βi S(x)i ,
αi + βi = 1,
where T and S are the temporal and spatial paths defined in Section 2.2. Each layer computes a gate for every token on every forward pass. Because xi includes positional encoding, weights can vary with content, channel–time position, layer, and input sample. AXON-TokenGated uses one full-length temporal window and no global path; final AXON replaces T with the two-window path (Section 2.5). Gate interventions are in Section 3.2 and Appendix H. 2.4
Global path
A natural extension adds a third path: dense attention over all tokens, with its own projections. The gate then predicts three weights that sum to one: yi = αi T (x)i + βi S(x)i + γi G(x)i , 3
αi + βi + γi = 1,
Table 1: Main EEG results — Linear Probe (LP, frozen encoder) and Full Finetune (FT). Balanced accuracy, mean ± std over 3 downstream seeds; subject-disjoint splits shared across all models. Model Dense Dense-L Divided-ST AXON-Fixed AXON-TokenGated AXON
motor
workload
hmc
siena
adftd
bcic
Mean LP
0.353±.002 0.384±.003 0.393±.001 0.370±.003 0.447±.001 0.455±.005
0.612±.021 0.648±.011 0.621±.008 0.674±.007 0.662±.012 0.673±.009
0.660±.003 0.653±.002 0.649±.002 0.666±.000 0.673±.001 0.650±.002
0.876±.007 0.828±.000 0.872±.003 0.841±.001 0.855±.000 0.867±.001
0.516±.034 0.531±.004 0.523±.028 0.502±.008 0.531±.014 0.538±.006
0.282±.008 0.288±.001 0.299±.002 0.296±.003 0.298±.004 0.292±.007
0.550 0.555 0.560 0.558 0.578 0.579
Full Finetune Model Dense Dense-L Divided-ST AXON-Fixed AXON-TokenGated AXON
motor
workload
hmc
siena
adftd
bcic
Mean FT
0.561±.015 0.623±.001 0.600±.006 0.619±.009 0.630±.012 0.625±.005
0.648±.008 0.613±.012 0.661±.057 0.735±.008 0.670±.033 0.685±.006
0.736±.003 0.728±.005 0.720±.003 0.731±.006 0.729±.003 0.732±.009
0.853±.001 0.826±.009 0.853±.011 0.853±.013 0.867±.013 0.859±.007
0.504±.046 0.536±.011 0.555±.021 0.501±.029 0.563±.005 0.604±.022
0.335±.038 0.469±.051 0.379±.017 0.403±.020 0.399±.046 0.428±.033
0.606 0.633 0.628 0.640 0.643 0.656
P where G(x)i = j∈Ω gij VG xj . This is AXON-withGlobal. The motivation was that clinical tasks may benefit from direct all-to-all context in a single layer, which the two axis paths only reach after two layers. AXON does not use this path; Section 3.2 reports its effect. 2.5
Two temporal windows
EEG carries information at different time scales: a motor-imagery response or a seizure onset develops over a few seconds, while a sleep stage or a cognitive state lasts the whole window [Pfurtscheller and Lopes da Silva, 1999]. A temporal path that always sees the whole window can in principle use both, but it has to learn to separate a short event from the slow background on its own. We make the separation explicit by giving the temporal path two windows: a short one that sees only five consecutive patches, and a long one that sees the whole visible window. Each layer learns how much to use each: Tms (x)i = λs Tshort (x)i + (1 − λs ) Tlong (x)i , λs ∈ [0, 1], with one λs per layer, initialised to favour the long window so that early reconstruction is easy. This choice is separate from the axis gate: the axis gate sets the temporal/spatial split for each token, and λs sets how far in time the temporal path looks. AXON, the final model, is AXON-TokenGated with this two-window temporal path and no global path.
3
Experiments
3.1
Setup
All encoders are pretrained with masked autoencoding [He et al., 2022] on pooled TUH-EEG [Obeid and Picone, 2016], I-CARE [Amorim et al., 2023], and internal EEG data for 50 epochs with batch size 4096. Recordings are mapped to 21 canonical 10–20 electrode positions and divided into 10second windows with T = 11 patches per channel. We mask 55% of tokens; the encoder processes only visible tokens, and a small decoder reconstructs masked patches. Missing channels are excluded from tokenization and reconstruction loss; attention masks cover only present tokens. We evaluate on six public datasets covering motor imagery (motor, bcic), cognitive workload (workload), sleep staging (hmc), seizure detection (siena), and dementia diagnosis (adftd). We report balanced accuracy (BAC) on subject-disjoint splits using linear probing (LP; shallow MLP head) and full fine-tuning (FT). For each model, Table 1 reports the checkpoint selected by the highest mean LP on held-out validation subjects, rather than the final epoch. Dataset summaries appear in Table 8; implementation, training-budget analysis, and evaluation details are in Appendices C and D. 3.2
Results
AXON achieves the highest Mean LP (0.579) and Mean FT (0.656) in Table 1. It improves LP and FT over Dense by 2.9 and 5.0 balanced-accuracy points, respectively. The gains remain over Dense-L 4
Table 2: Controlled audio spectrogram results — full AudioSet-2M pretraining. All variants use identical optimizer, schedule, and compute budget; only the attention operator differs. Model Dense Fixed gate (AXON-Fixed) Token gate (AXON-TokenGated) Token + Global (AXON-withGlobal)
AudioSet FT mAP
AudioSet LP mAP
ESC-50 LP Acc
SC LP Acc
11.59±0.25 13.75±0.16 13.82±0.25 14.29±0.05
4.08±0.01 4.01±0.02 4.84±0.03 5.08±0.04
45.25±0.94 44.83±0.31 45.75±0.71 47.17±0.92
19.53±0.07 22.01±0.06 21.96±0.15 21.87±0.16
(2.4 and 2.3 points) and Divided-ST (1.9 and 2.8 points). The largest task-level gains over Dense are on motor LP (+10.2 points) and adftd FT (+10.0 points). Three further controls show that the gain is not simply from attending to fewer tokens, and not simply from position information: restricting the spatial path to each electrode’s nearest neighbours drops Mean LP to 0.510; the trained dense model still attends to pairs that share neither electrode nor time step (Table 16); and adding 2D rotary position embeddings [Su et al., 2024] to the dense model gains only +0.008 (Appendix F). The ranking is stable at every pretraining budget we tested (Appendix D.2). A variant that biases the temporal path toward nearby time steps (AXON-Decay) matches AXON on Mean LP but lowers Mean FT (Appendix F.1). Appendix H shows what the gate learns and tests whether the model depends on it. AXON-withGlobal, the variant with a third dense path (Section 2.4), underperforms both AXON and AXON-Fixed on Mean LP and FT (Table 11) despite reaching comparable reconstruction loss (Figure 2). Its first-layer global weight is γ = 0.724, leaving about a quarter for the axis paths. We hypothesise that global averaging provides a reconstruction shortcut that weakens the learning of axis-specific features (Appendix G). Comparison with released EEG foundation models. Under the same six-task protocol, AXON has the highest mean scores among six released EEG foundation models: 0.579 vs. 0.527 Mean LP and 0.656 vs. 0.614 Mean FT against the next-best, REVE (Appendix B).
4
Audio Spectrogram Experiments
We apply the same axis-factorization principle to time–frequency spectrograms, comparing four attention variants under the AudioMAE pretraining recipe [Huang et al., 2022]. A spectrogram is a grid too: one axis is time, the other is frequency, and a token is one frequency band over one time frame. The temporal path attends across time within a frequency band; the spatial path becomes a frequency path and attends across frequency bands within a time frame. The full-scale experiment uses AudioSet-2M corpus [Gemmeke et al., 2017] (∼2M clips); smaller-scale results and the comparison with published AudioMAE are in Appendix J. Downstream evaluation uses AudioSet, ESC-50 [Piczak, 2015] and SpeechCommands (SC) [Warden, 2018]; AudioSet is multi-label, so we report mean average precision (mAP), and the other two report accuracy. All factorized variants improve over Dense on AudioSet FT and SpeechCommands LP (Table 2). Unlike EEG, Token+Global achieves the highest scores on three of the four metrics. On AudioSet FT and SpeechCommands LP this advantage holds at all three pretraining scales we tested, 18K, 200K and 2M clips (Tables 18 and 19). The audio experiment therefore extends the factorization result while showing that the global branch is not uniformly detrimental across the tested settings.
5
Conclusion
AXON combines temporal and spatial attention through learned soft mixing. Across six subjectdisjoint EEG tasks, it improves mean balanced accuracy over dense and parameter-matched dense baselines under Linear Probing and full fine-tuning. Gate interventions show larger performance drops from removing either axis or using hard routing than from replacing token gates with layer means. Axis factorization extends to audio, a second axis-structured modality. These results support a simple design rule for structured signals: align attention with the signal’s axes and let the model learn how much to use each.
5
References Diego Alvarez-Estevez and Roselyne Rijsman. Haaglanden Medisch Centrum sleep staging database. PhysioNet, 2022. Version 1.1. Edilberto Amorim, Wei-Long Zheng, Mohammad Ghassemi, Mahsa Aghaeeaval, Pradyot Kandhare, Vishal Karukonda, Jong Woo Lee, Susan T. Herman, Adithya Sivaraju, Nicolas Gaspard, Jeannette Hofmeijer, Michel J. A. M. van Putten, Reza Sameni, Matthew A. Reyna, Gari D. Clifford, and M. Brandon Westover. The International Cardiac Arrest Research Consortium electroencephalography database. Critical Care Medicine, 2023. doi: 10.1097/CCM.0000000000006074. Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. ViViT: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021. Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15750–15758, 2021. Paolo Detti. Siena scalp EEG database. PhysioNet, 2020. Version 1.0.0. Yassine El Ouahidi, Jonathan Lys, Philipp Thölke, Nicolas Farrugia, Bastien Pasdeloup, Vincent Gripon, Karim Jerbi, and Giulia Lioi. REVE: A foundation model for EEG – adapting to any setup with large-scale pretraining on 25,000 subjects. In Advances in Neural Information Processing Systems 38, 2025. arXiv:2510.21585. Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 776–780, 2017. doi: 10.1109/ICASSP.2017.7952261. Ary L. Goldberger, Luı́s A. N. Amaral, Leon Glass, Jeffrey M. Hausdorff, Plamen Ch. Ivanov, Roger G. Mark, Joseph E. Mietus, George B. Moody, Chung-Kang Peng, and H. Eugene Stanley. PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation, 101(23):e215–e220, 2000. doi: 10.1161/01.CIR.101.23.e215. Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16000–16009, 2022. Byeongho Heo, Song Park, Dongyoon Han, and Sangdoo Yun. Rotary position embedding for vision transformer. In European Conference on Computer Vision (ECCV), 2024. arXiv:2403.13298. Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180, 2019. Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer. Masked autoencoders that listen. In Advances in Neural Information Processing Systems 35, 2022. Herbert H. Jasper. The ten-twenty electrode system of the International Federation. Electroencephalography and Clinical Neurophysiology, 10:371–375, 1958. Weibang Jiang, Liming Zhao, and Bao-liang Lu. Large brain model for learning generic representations with tremendous EEG data in BCI. In The Twelfth International Conference on Learning Representations, 2024. Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network representations revisited. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pages 3519–3529, 2019. 6
Wei Lun Lim, Olga Sourina, and Lipo Wang. STEW: Simultaneous task EEG workload dataset. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 26(11):2106–2114, 2018. doi: 10.1109/TNSRE.2018.2872924. Andreas Miltiadous, Katerina D. Tzimourta, Theodora Afrantou, Panagiotis Ioannidis, Nikolaos Grigoriadis, Dimitrios G. Tsalikakis, Pantelis Angelidis, Markos G. Tsipouras, Euripidis Glavas, Nikolaos Giannakeas, and Alexandros T. Tzallas. A dataset of scalp EEG recordings of Alzheimer’s Disease, Frontotemporal Dementia and healthy subjects from routine EEG. Data, 8(6):95, 2023. doi: 10.3390/data8060095. Iyad Obeid and Joseph Picone. The Temple University Hospital EEG data corpus. Frontiers in Neuroscience, 10:196, 2016. doi: 10.3389/fnins.2016.00196. Gert Pfurtscheller and F. H. Lopes da Silva. Event-related EEG/MEG synchronization and desynchronization: basic principles. Clinical Neurophysiology, 110(11):1842–1857, 1999. doi: 10.1016/S1388-2457(99)00141-8. Karol J. Piczak. ESC: Dataset for environmental sound classification. In Proceedings of the 23rd ACM International Conference on Multimedia, pages 1015–1018, 2015. doi: 10.1145/2733373.2806390. Olivier Roy and Martin Vetterli. The effective rank: A measure of effective dimensionality. In 2007 15th European Signal Processing Conference (EUSIPCO), pages 606–610, 2007. Gerwin Schalk, Dennis J. McFarland, Thilo Hinterberger, Niels Birbaumer, and Jonathan R. Wolpaw. BCI2000: A general-purpose brain-computer interface (BCI) system. IEEE Transactions on Biomedical Engineering, 51(6):1034–1043, 2004. doi: 10.1109/TBME.2004.827072. Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. RoFormer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. doi: 10.1016/j.neucom.2023.127063. arXiv:2104.09864 (2021). Michael Tangermann, Klaus-Robert Müller, Ad Aertsen, Niels Birbaumer, Christoph Braun, Clemens Brunner, Robert Leeb, Carsten Mehring, Kai J. Miller, Gernot R. Müller-Putz, Guido Nolte, Gert Pfurtscheller, Hubert Preissl, Gerwin Schalk, Alois Schlögl, Carmen Vidaurre, Stephan Waldert, and Benjamin Blankertz. Review of the BCI Competition IV. Frontiers in Neuroscience, 6:55, 2012. doi: 10.3389/fnins.2012.00055. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems 30, 2017. Guangyu Wang, Wenchao Liu, Yuhong He, Cong Xu, Lin Ma, and Haifeng Li. EEGPT: Pretrained transformer for universal and reliable representation of EEG signals. In Advances in Neural Information Processing Systems 37, pages 39249–39280, 2024. Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Tao Li, and Gang Pan. CBraMod: A criss-cross brain foundation model for EEG decoding. In The Thirteenth International Conference on Learning Representations, 2025. Pete Warden. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209, 2018. Chaoqi Yang, M. Brandon Westover, and Jimeng Sun. BIOT: Biosignal transformer for crossdata learning in the wild. In Advances in Neural Information Processing Systems 36, pages 78240–78260, 2023. Yuchen Zhou, Jiamin Wu, Zichen Ren, Zhouheng Yao, Weiheng Lu, Kunyu Peng, Qihao Zheng, Chunfeng Song, Wanli Ouyang, and Chao Gou. CSBrain: A cross-scale spatiotemporal brain foundation model for EEG decoding. In Advances in Neural Information Processing Systems 38, 2025. arXiv:2506.23075.
7
Appendix A
Background & mannas.ai Context
Electroencephalography (EEG) measures electrical potential differences across a sparse array of scalp electrodes at millisecond resolution. Each recording is a matrix of C channels by T time samples. The signal is low-amplitude (∼10–100 µV), contaminated by muscle artifacts, eye movements, and ambient noise, and varies across subjects and sessions. Despite this noise, EEG encodes clinically and cognitively meaningful structure: motor imagery induces localized mu-rhythm desynchronization over sensorimotor cortex; sleep staging depends on broadband slow-wave and spindle patterns; epileptic events manifest as sharp, spatially propagating discharges. Two properties make EEG a natural proving ground for anisotropic attention. First, the token mannas.ai is explicitly two-dimensional: a patch token at (c, t) has a known spatial identity (electrode c with head coordinate) and a temporal identity (time window t). Second, different tasks depend on structurally different contexts: motor imagery requires preserving localized, lateralized, temporally precise structure; clinical classification benefits from broader spatial and temporal integration. A single isotropic attention operator may be a poor shared prior across tasks with such different channel-time structure.
B
External EEG Foundation Models
B.1
External EEG baselines — Linear Probe
We compare AXON against six published EEG foundation models under a standardised evaluation protocol: the same six tasks, the same disjoint subject splits, the same balanced-accuracy metric, and the same three downstream seeds as our internal ablations. All external models are evaluated from their officially released pretrained checkpoints. Models are evaluated with both linear probing (LP, frozen encoder) and fine-tuning (FT, all weights updated).
Table 3: External EEG foundation model comparison — Linear Probe (LP, frozen encoder). Format: balanced accuracy. Bold = column best. Our dense baseline is included as a within-setup reference. All models evaluated under the same 6-task protocol with identical subject splits. Model
motor
workload
hmc
siena
adftd
bcic
Mean
EEGPT [Wang et al., 2024] LaBraM [Jiang et al., 2024] CBraMod [Wang et al., 2025] BIOT [Yang et al., 2023] CSBrain [Zhou et al., 2025] REVE [El Ouahidi et al., 2025]
0.381±.015 0.268±.012 0.259±.013 0.284±.008 0.272±.005 0.315±.003
0.574±.020 0.500±.000 0.500±.000 0.577±.081 0.500±.000 0.709±.035
0.665±.006 0.381±.019 0.510±.001 0.644±.003 0.568±.002 0.653±.006
0.804±.019 0.500±.000 0.619±.006 0.610±.025 0.500±.000 0.688±.033
0.393±.034 0.309±.019 0.358±.002 0.492±.033 0.364±.014 0.532±.052
0.281±.007 0.285±.022 0.270±.010 0.261±.010 0.270±.009 0.267±.012
0.516 0.374 0.419 0.478 0.412 0.527
Dense AXON
0.353±.002 0.455±.005
0.612±.021 0.673±.009
0.660±.003 0.650±.002
0.876±.007 0.867±.001
0.516±.034 0.538±.006
0.282±.008 0.292±.007
0.550 0.579
AXON achieves the highest Mean LP and Mean FT among all eight models, ahead of the next-best external model (REVE) by +5.2 points LP (+9.9% relative) and +4.2 points FT (+6.8% relative). The full fine-tuning comparison is in Appendix Table 4. Note that our own dense baseline already outperforms all six external models on mean LP. So part of AXON’s gap to the external models comes from our training setup rather than from the attention design, and AXON’s gain over that dense baseline (Table 1) is measured on top of it. The architectural claim rests on that controlled comparison, not on this table. Task-level analysis reveals that motor imagery exhibits the largest gap between AXON and external models (+19% LP over the best competitor). This fits the motivation for preserving axis structure: motor imagery depends on a left/right difference in mu and beta rhythms over the sensorimotor electrodes. This comparison does not, however, show which features account for the gain. REVE retains advantages on workload and siena FT, suggesting its architecture provides broader temporal integration suited for sustained cognitive states. 8
Table 4: External EEG foundation model comparison — Full Finetune (FT). Bold = column best. All models evaluated under the same 6-task protocol with identical subject splits. Model
motor
workload
hmc
siena
adftd
bcic
Mean
EEGPT [Wang et al., 2024] LaBraM [Jiang et al., 2024] CBraMod [Wang et al., 2025] BIOT [Yang et al., 2023] CSBrain [Zhou et al., 2025] REVE [El Ouahidi et al., 2025]
0.513±.007 0.249±.001 0.426±.015 0.372±.013 0.570±.023 0.612±.004
0.668±.015 0.501±.002 0.554±.044 0.570±.092 0.605±.028 0.703±.011
0.712±.005 0.645±.011 0.709±.011 0.710±.005 0.705±.010 0.724±.002
0.795±.031 0.813±.018 0.847±.036 0.747±.026 0.782±.036 0.863±.033
0.404±.022 0.290±.060 0.319±.026 0.449±.038 0.432±.021 0.460±.050
0.280±.041 0.259±.009 0.303±.031 0.314±.040 0.350±.015 0.322±.014
0.562 0.460 0.526 0.527 0.574 0.614
Dense AXON
0.561±.015 0.625±.005
0.648±.008 0.685±.006
0.736±.003 0.732±.009
0.853±.001 0.859±.007
0.504±.046 0.604±.022
0.335±.038 0.428±.033
0.606 0.656
B.2
External EEG baselines — Full Finetune
C
Extended Implementation Details & Pretraining
Pretraining corpus. The pretraining corpus pools four data sources (Table 5). Two are publicly available: the Temple University Hospital EEG Corpus (TUH-EEG) [Obeid and Picone, 2016] and the I-CARE dataset [Amorim et al., 2023]. Two are internal clinical EEG collections acquired under institutional ethics approval and de-identified before use; these are not publicly released but are described below to enable reproducibility assessment. Table 5: Pretraining corpus composition. All sources are pooled into a single unlabeled pretraining set; no downstream task labels are used during pretraining. Internal sources are marked †. Source
Subjects
Hours
TUH-EEG [Obeid and Picone, 2016] I-CARE [Amorim et al., 2023] Internal-A † Internal-B †
∼15,000 ∼600 4,539 1,050
∼25,000 ∼33,000 2,546 435
Clinical context
Availability
Mixed clinical referrals Post-cardiac-arrest ICU Routine clinical neurophysiology Multi-centre research EEG
Public Public Not released Not released
Internal-A comprises routine clinical EEG recordings (resting-state, hyperventilation, and photic stimulation protocols) collected across hospital neurophysiology departments. Internal-B comprises multi-centre research EEG recordings acquired under a national research programme. Both internal datasets were recorded with standard 10-20 montage systems at sampling rates of 250–512 Hz and deidentified (all patient identifiers, dates, and institution codes removed) before inclusion. All internal data collection was conducted under institutional ethics board approval with informed consent or waiver of consent for retrospective de-identified use. Subject identities are verified disjoint across all four pretraining sources and all six downstream evaluation datasets. Recordings span diverse acquisition settings with variable electrode configurations: systems range from compact 16-channel ambulatory devices to full 256-channel research amplifiers, and not every recording contains all standard 10-20 electrodes. We retain only recordings whose channel header resolves to a subset of the international 10-20 montage, then extract the available 10-20 electrode positions per recording. Because channel count varies across sources, we adopt C = 21 as the representative value throughout this paper; this is the mode channel count across the retained pretraining corpus. Preprocessing: (1) resample to fs = 200 Hz; (2) notch filter at 50 and 60 Hz; (3) bandpass [0.5, 99.5] Hz; (4) per-channel z-score normalisation; (5) clip at ±15σ; (6) segment into non-overlapping 10-second windows. We train encoders using masked autoencoding (MAE). EEG signal is divided into patches of 200 samples (1 s at fs = 200 Hz) with a 20-sample overlap and 180-sample (0.9 s) stride between consecutive patch start positions, yielding T = 11 patches per channel per 10-second window. Each channel-time patch becomes a single token via a learnable linear projection. Tokens receive a split positional encoding: spatial coordinates pc = (x, y, z) (3D head positions in millimetres, standard 10-20 montage) are projected with a learned linear layer to produce PES ; the temporal patch index receives a fixed sinusoidal encoding PET . The two components are summed: PE(i) = PES (ci ) + PET (ti ). The encoder processes only the visible tokens (55% masking ratio, spatiotemporal block masking with spatial radius 3.0 and temporal radius 3.0); masked token positions receive no encoder gradient. A lightweight 4-layer dense Transformer decoder takes the encoded visible tokens plus 9
learned mask-slot embeddings and reconstructs all patches. Training minimises L1 reconstruction loss over masked patches plus an auxiliary pooled-attention reconstruction loss (λ = 0.5). The auxiliary head applies cross-attention pooling over the concatenated outputs of all encoder MHA layers: a single learned query token attends over the layer-wise output tokens to produce a compact global representation. This pooled token is then repeated to match the number of masked positions, enriched with positional encodings, and passed through a 2-layer FFN to reconstruct the masked patches under a separate L1 loss. The total pretraining loss is L = Lprimary + λ · Laux . C.1
Pretraining hyperparameters
Table 6: Pretraining hyperparameters (AXON). The dense baseline uses the same schedule and corpus; it differs only in the encoder attention operator. Hyperparameter
Value
Batch size Epochs Peak LR LR schedule Optimiser Gradient clip (ℓ2 norm) Gate τ warmup Mask ratio Spatial block radius Temporal block radius Dropout mask ratio Coordinate noise σ Auxiliary loss weight Model dim d Encoder layers / heads Decoder layers Precision
4096 50 2.4×10−4 CosineAnnealingLR (Tmax = 20); fused AdamW (β1 = 0.9, β2 = 0.95, λ = 0.05) 1.0 2.0 → 1.0 over 1500 steps 0.55 (spatiotemporal block masking) 3.0 patch indices (i.e., 3 electrode positions in the canonical 10-20 ordering) 3.0 patch indices (i.e., 3 consecutive time patches, ≈2.7 s) 0.3 0.25 0.5 512 22 / 8 4 (dense, full attention) bfloat16 (autocast)
Pretraining was conducted on a single compute node equipped with 8 NVIDIA H200 GPUs (143,771 MiB memory each, ≈1.1 TiB total GPU memory) and ≈2.2 TiB of host RAM. C.2
Parameter count and compute
AXON is not parameter-matched to the dense baseline. Replacing one dense attention operator with two separate full-width temporal and spatial operators adds approximately 26M parameters (22.5%). Independent projections. The temporal and spatial paths use separate QKV and output projection matrices (Q(T ) , K (T ) , V (T ) , O(T ) ) and (Q(S) , K (S) , V (S) , O(S) ). This is deliberate: the features needed to select which temporal patch to attend to (e.g., spectral power at a rhythm band) are different from those needed to select which electrode to attend to (e.g., lateralised activation). Shared projections would force both axes through the same feature bottleneck, limiting specialisation. The cost is a 2× parameter increase in the attention projections per layer relative to a single dense attention, but the number of pairwise attention interactions is reduced by a factor of ≈ (CT )/(C + T ) (from (CT )2 to CT 2 + T C 2 ). To control for this, we trained a parameter-matched dense baseline with hidden dimension 568 (142.5M total parameters, within 0.5% of AXON’s 141.9M). This model uses identical pretraining (same data, optimizer, epochs, mask ratio) and differs only in hidden dimension. Three observations address the parameter concern. First, the parameter-matched dense baseline (dim=568, 142.5M) uses nearly identical capacity to AXON (141.9M) with the same dense attention topology. It achieves Mean LP 0.555 and Mean FT 0.633: higher than the original dense baseline (0.550 / 0.606), demonstrating that the extra parameters provide some benefit, but still falling short of AXON by 2.4 points on Mean LP and 2.3 points on Mean FT (+4.3% and +3.6% relative). The AXON advantage persists after capacity matching. Second, AXON replaces global quadratic mixing with two axis-factorized operators, reducing pairwise attention interactions by ∼7× per layer despite 10
Table 7: Parameter count, attention complexity, and downstream performance. Dense-L controls for the capacity difference by matching AXON’s total parameter count. Model Dense Dense-L AXON-withGlobal AXON ‡
Total params
Attn ops/layer‡
Mean LP
Mean FT
115.9M 142.5M ∼167M 141.9M
O((CT )2 ) = 53,361 O((CT )2 ) = 53,361 > (CT )2 O(CT 2 + T C 2 ) = 7,392
0.550 0.555 0.555 0.579
0.606 0.633 0.619 0.656
Computed for C = 21 channels, T = 11 time patches (mode of pretraining corpus): Dense (CT )2 = 53,361; AXON CT 2 + T C 2 = 7,392 (≈7× less).
the added projections. Because the sequence length (N = 231) is small relative to d = 512, the O(N d2 ) projection costs dominate total FLOPs; the computational advantage of axis factorization therefore lies in the structured prior, not in raw speed. Third, the AXON-withGlobal variant has substantially more parameters (∼167M, adding a global QKV on top of temporal and spatial branches) and still underperforms AXON by 2.4 points on Mean LP and 3.7 points on Mean FT. If the gains were explained by parameter count, AXON-withGlobal should win. The parameter-matched dense and AXON-withGlobal comparisons together isolate the inductive bias, not a capacity effect. Empirical gate statistics. The gate diagnostic (22 layers × 6 datasets) shows that the trained gate is neither trivially uniform nor collapsed. Mean axis weights: αmean = 0.441 (temporal), βmean = 0.559 (spatial). Mean gate entropy ratio: 0.932 (scale 0–log 2, where 1.0 is fully uniform). No layer falls below the collapse threshold of 0.40. Gate intervention experiments (Table 14) show that the dominant learned structure is a soft depth-dependent anisotropy schedule (layer-mean gates drop only −2.1% vs learned), while exact per-token gate assignment is a secondary effect (shuffled gates drop only −1.3%). Stop-gradient. The stop-gradient on the gate input means the encoder receives no gradient from the routing decision; it is trained only by the reconstruction loss. We included it as a precaution so that the gate could not reshape the encoder’s representations during training. The ablation shows it makes no measurable difference: removing it, so that the gate reads the live representation instead of a detached copy [Chen and He, 2021], changes Mean LP by only −0.002 (0.577 vs. 0.579). It is a safe default, not a source of gain; our reported results do not hinge on it. Temperature annealing. The gate softmax is divided by a temperature τ annealed from τstart = 2.0 to τend = 1.0 over the first 1500 training steps: t τt = τstart + (τend − τstart ) · min 1, . 1500 During warmup the gate is soft, allowing both axes to receive gradient from the MAE reconstruction loss. Both branches therefore develop useful representations before the gate sharpens. Gate entropy is monitored throughout warmup; a drop below 0.40 before warmup ends would indicate premature routing commitment and would be corrected by increasing τstart or extending the warmup window.
D
Downstream Tasks & Evaluation Protocol
We evaluate with two protocols: linear probing (LP), where the encoder is frozen and only the classification head is trained, and full finetuning (FT), where all encoder weights are updated. The classification head is: AdaptiveAvgPool1d → Linear(512, 128) → ELU → Dropout(0.3) → Linear(128, K), where K is the number of classes. The downstream optimiser is AdamW (weight decay 0.01) with cosine-annealing LR (peak 2×10−4 , min 2×10−5 ) preceded by a 5-epoch linear warmup from 2×10−6 . 30 training epochs. The metric is balanced accuracy throughout. Subject splits are disjoint across train, validation, and test for every dataset; the same splits are shared by all models. 11
Table 8: Downstream evaluation datasets. All splits are subject-disjoint. Dataset
Task
Classes
motor mv img [Schalk et al., 2004, Goldberger et al., 2000] bcic 2a [Tangermann et al., 2012] workload [Lim et al., 2018] hmc [Alvarez-Estevez and Rijsman, 2022] siena scalp [Detti, 2020] adftd [Miltiadous et al., 2023]
Motor imagery (L/R/both/foot) Motor imagery BCI (4 limbs) Cognitive workload (low/high) Sleep staging (5 stages) Seizure detection (ictal/interictal) Dementia (AD/FTD/Healthy)
4 4 2 5 2 3
D.1
Window 4s 4s 4s 30 s 10 s 10 s
Train split
Val / Test split
Subj. 0–69 Subj. 1–5 ≈72% subj. ≈67.6% subj. ≈70% subj. ≈70% subj.
Subj. 70–88 / 89–109 Subj. 6–7 / 8–9 ≈14% / 14% subj. ≈16.2% / 16.2% subj. ≈15% / 15% subj. ≈15% / 15% subj.
Downstream tasks
We evaluate across six datasets covering BCI, cognitive, and clinical tasks. All evaluations use disjoint subject splits and balanced accuracy. Table 9 summarises the key discriminative challenge of each task and explains why isotropic attention is an unfavourable inductive bias. Table 9: Downstream evaluation tasks. Dataset
Family
Classes
motor mv img bcic 2a workload hmc siena scalp adftd
BCI BCI Cognitive Clinical Clinical Clinical
4 4 2 5 2 3
Key discriminative structure Lateralized mu/beta ERD at movement onset Fine lateralization, 4 limb classes Sustained frontal-parietal synchrony Broadband spectral stage transitions Spatially propagating ictal discharge Diffuse cortical slowing, theta excess
Motor imagery tasks are most sensitive to interaction-isotropic mixing: mu/beta ERD lateralization [Pfurtscheller and Lopes da Silva, 1999] (left vs. right hand) is the primary discriminative signal, and averaging across all tokens via dense global attention can suppress this asymmetry. D.2
Training budget
Table 1 reports the best-validation-Mean-LP checkpoint (held-out subjects, never test), not the final epoch. To show the budget does not drive the result, we linear-probed AXON and Dense at epochs 5–50 (Mean LP over the six tasks, single seed; Table 10). Validation-best was epoch 10 (AXON) and epoch 9 (Dense); at the final epoch AXON still leads 0.568 vs. 0.543. The ranking is stable at every budget: AXON leads by +0.025–0.028 Mean LP and Dense never catches up, so the gain is not faster learning that more compute would erase. We do not claim it holds for unlimited training, only that it is stable across every budget we tested. Table 10: Mean LP over the six tasks at matched pretraining budgets (single seed). AXON leads at every epoch. Epoch
5
10
20
30
40
50
AXON Dense
0.584 0.556
0.579 0.551
0.580 0.552
0.574 0.549
0.570 0.545
0.568 0.543
E
Why Two Layers Connect Every Pair of Tokens
E.1
Axis graph diameter
AXON removes the dense all-to-all path. In one layer a token attends only to tokens on its own electrode (the temporal path) and to tokens at its own time step (the spatial path). A concern is that this cuts the model off from the rest of the recording: a token on electrode c at time t never sees electrode c′ at time t′ directly. The proposition below shows that this is not so. On a full channel-time grid, any two tokens are connected after two layers, so the model keeps its global reach. What changes is which pairs interact directly within one layer, and how many. This is why we describe the design as changing which pairs interact directly, not the model’s reach, and it is the fact behind the statement in Section 2.2 that any two tokens can still meet after two layers. 12
Proposition 1 (Axis graph has diameter at most two). On the full channel-time grid Ω = C × T , the graph with edges between tokens that share either the same channel or the same time index has diameter at most two. After two stacked axis-attention layers, any token can receive information from any other token. The number of possible one-hop attention interactions is |Eaxis | ≤ CT 2 + T C 2 = CT (C + T ), versus |Edense | = (CT )2 for the dense graph. Proof. Take any two tokens u = (c, t) and v = (c′ , t′ ). If c = c′ or t = t′ , then u and v are connected by one axis edge. Otherwise, u is connected to (c′ , t) by a spatial edge (same time step), and (c′ , t) is connected to v = (c′ , t′ ) by a temporal edge (same channel). Every pair is thus connected by a path of length at most two. The edge-count bound follows from C temporal groups of size T and T spatial groups of size C.
F
Ablation Summary
F.1
Interpreted ablation summary
Table 11: EEG ablation summary. Values are unweighted mean balanced accuracy across six downstream tasks. Diagnostic variants are included to show task-specific tradeoffs; bold indicates the best mean in each column among evaluated variants. Variant
Mean LP
Mean FT
Dense
0.550
0.606
Dense-L Divided-ST
0.555 0.560
0.633 0.628
Dense + joint PE
0.540
0.600
Dense + 2D-RoPE
0.554
—
AXON-Fixed
0.558
0.640
Fixed gate, schedule init
0.555
0.636
AXON-TokenGated
0.578
0.643
AXON
0.579
0.656
AXON without stop-gradient
0.577
—
AXON-withGlobal
0.555
0.619
KNN spatial constraint
0.510
0.609
Head-split multiscale
0.553
0.628
Temporal decay
0.556
0.642
AXON-Decay continuation
0.579
0.633
Purpose of comparison Dense attention baseline with split positional encoding. Parameter-matched dense control. Sequential divided space-time operator [Bertasius et al., 2021] in the identical setup (Appendix F.2). Tests whether joint positional encoding improves over split PE. Axial 2D rotary position embedding on the dense graph; single seed vs. a single-seed dense reference (0.546). Tests axis factorization without token-dependent routing. Tests whether initializing a fixed gate to AXON’s learned layer schedule is sufficient. Tests token-gated temporal/spatial factorization without multiscale temporal windows. Final no-global model with token axis gate and twoscale temporal branch. Gate reads the live representation instead of a detached copy; single seed. Tests whether adding a dense global branch improves the factorized encoder. Tests whether local electrode-neighbour masking helps the spatial path. Tests an alternative multiscale temporal implementation. Diagnostic locality bias; helps some event-like tasks but reduces mean performance. Diagnostic continuation run; maintains LP but reduces FT.
Reading the ablations. The ablation summary supports three conclusions. First, axis factorization improves over dense attention even after controlling for parameter count (Dense-L). Second, token gating provides the largest additional LP gain over fixed factorization (+0.020), while the two-scale temporal path contributes a smaller task-selective refinement. Third, the negative variants show that adding a dense global branch, hard spatial masks, or extra temporal-scale machinery does not improve mean transfer. AXON is the only variant that achieves both the best mean LP and the best mean 13
Table 12: Architectural variants summary. All models use the same pretraining corpus and schedule; they differ only in attention structure. Variant
Temporal windows
Global path
Key change
Mean LP
Dense Divided-ST AXON-Fixed AXON-TokenGated AXON-withGlobal AXON AXON-Decay
W=−1 W=−1 W=−1 W=−1 W=[5,−1] [5,−1] [5,−1] + decay
Full No No No Yes No No
Reference Sequential T→S blocks Fixed factorized axes Per-token axis gating +Global dense path +Two-scale gate +Temporal locality
0.550 0.560 0.558 0.578 0.555 0.579 0.579
Head-split multiscale 3-scale gate
[5,−1] static [2,5,−1]
No No
Static scale split 3-scale gate
0.553 0.558
FT; variants that improve selected tasks (e.g., AXON-Decay) do not improve the mean FT objective and are therefore treated as diagnostic, task-selective extensions. The token gate’s LP advantage over AXON-Fixed arises from training-time gradient diversity rather than inference-time routing; see Appendix H.4 for the schedule-initialized fixed gate experiment that isolates this effect. Three windows. We also tried three windows (two, five, and all patches) with an entropy penalty that pushes the scale mix to use all three. Mean LP fell to 0.558. The reconstruction loss barely distinguishes the three windows, so the penalty dominates and the mix stays close to uniform; two windows with a long-window start was the best we found. Temporal locality (AXON-Decay). AXON-Decay adds a learnable temporal-distance bias to the full-window temporal path, so that nearby time steps get more weight. It helps tasks driven by short events, motor imagery (LP 0.455 → 0.487) and seizure detection (siena LP 0.867 → 0.894); it matches AXON on Mean LP (0.579), lowers Mean FT (0.656 → 0.633), and hurts workload, a sustained-state task. So the best temporal context length depends on the task, and we treat AXON-Decay as a diagnostic, not a replacement for AXON. 2D RoPE control. We ran axial 2D RoPE [Su et al., 2024, Heo et al., 2024]: each head (dim 64) is split so attention depends only on the relative (∆t, ∆c) offset while the graph stays fully dense. Result (same pipeline as Table 1; single seed, vs. a single-seed dense reference): Dense 0.546 → Dense+2D-RoPE 0.554, a +0.008 Mean LP gain concentrated almost entirely on adftd (our highest-variance task, ±0.034 across seeds). Axis factorization gives +0.029, ≈3.6× the RoPE delta. Thus 2D RoPE does not substitute for factorization (though it does not fail either). This is expected: RoPE changes how position enters the scores but leaves the graph dense; every off-axis pair stays available, whereas AXON removes those edges by construction. After pretraining the dense encoder still places 30–63% of its attention on off-axis pairs (token pairs sharing neither the same electrode nor the same time step; Table 16, Figure 8), so it does not suppress those interactions on its own, and a relative-position code gives it no mechanism to. The two are orthogonal, not substitutes. F.2
Divided space-time baseline
The divided space-time baseline is pretrained in our exact EEG setup: same pooled corpus, same 10-20 montage and tokenisation, batch size 4096, peak LR 2.4 × 10−4 , fused AdamW, bfloat16, the identical pretraining budget, the identical MAE objective (L1 on masked patches plus the pooledattention auxiliary loss), and the same split positional encoding. Only the attention operator differs. We implement the divided space-time block of Bertasius et al. [2021] faithfully: attention along time, then attention along channels, each with its own residual connection and with the temporal output projection, in place of AXON’s parallel token-gated axis mixture. Parameter count and per-layer compute are comparable to AXON’s ∼141.9M. The two models converge to matched reconstruction loss, so neither is under-trained relative to the other. Metric is balanced accuracy on subject-disjoint splits identical across models; per-task results are in Table 1, where ± is the standard deviation across the 3 downstream seeds. 14
Divided-ST recovers part of the gap to AXON, so restricting attention to the two axes helps on its own. The rest of the gap is what the parallel, gated composition adds. In Divided-ST the axis order is fixed and identical for every token and every layer; in AXON the gate sets the temporal/spatial balance per token and per layer, and Appendix H shows that this learned per-layer balance is where most of the gate’s benefit lies.
G
Analysis of the Global-Path Variant
Section 2.4 states our hypothesis: under masked pretraining the dense global path is a shortcut, averaging over visible tokens to guess a masked patch, so the encoder builds fewer axis-specific features. Here we give the evidence behind that reading. Same reconstruction loss, different features. AXON and AXON-withGlobal reach the same reconstruction loss (Figure 2), yet AXON-withGlobal transfers worse (Mean LP 0.555 vs. 0.579, Mean FT 0.619 vs. 0.656). So the difference is in the features the two encoders build, not in how well they reconstruct.
Figure 2: MAE reconstruction loss during pretraining. AXON (no global path) and AXONwithGlobal converge to comparable reconstruction loss (≈0.38–0.39), yet AXON-withGlobal underperforms on mean LP (0.555 vs. 0.579) and mean FT (0.619 vs. 0.656). The global path does not improve reconstruction; it changes how the encoder reconstructs, routing mass through dense averaging (γ = 0.724 at layer 1) rather than through axis-structured features.
Effective rank. Effective rank [Roy and Vetterli, 2007] counts how many dimensions a set of features actually uses; a higher value means richer, less collapsed features. We compute it on the final mean-pooled embeddings of up to 512 validation samples per dataset (Table 13). Removing the global path raises the effective rank on five of the six datasets. The largest jump is on motor imagery, from 19.97 to 29.05, and motor imagery is also the task where AXON gains most over Dense (Table 1). This fits the shortcut reading: motor imagery depends on fine, local spatio-temporal structure that averaging over all tokens washes out. CKA similarity. Linear CKA [Kornblith et al., 2019] measures how similar two sets of features are (1.0 identical, 0.0 unrelated). The features of AXON and AXON-withGlobal are least similar on motor imagery (0.84), sleep staging (0.86) and workload (0.87), and almost identical on seizure detection and dementia (above 0.96). The global path changes the features most on the tasks with the most temporal or spatial structure, and least where the two models also perform alike. Per-layer curves for all six datasets are in Figure 3. 15
Table 13: Effective rank and representation similarity (Linear CKA) of final mean-pooled embeddings between the factorized noGlobal model and the withGlobal model. Dataset Effective Rank ↑ CKA Similarity ↓
withGlobal
noGlobal
withGlobal vs noGlobal
17.76 66.08 22.83 19.97 28.93 51.58
21.29 77.66 28.79 29.05 27.19 65.48
0.9748 0.9363 0.8662 0.8412 0.9634 0.8698
adftd bcic 2a hmc motor siena workload
(a) Layer-wise effective rank, all six datasets.
(b) Layer-wise Linear CKA, all six datasets.
Figure 3: Layer-wise effective rank and CKA divergence between factorized and dense-global representations.
H
Gate Mechanism Analysis
H.1
Understanding the gate
Figure 4 shows the average gate weight per layer over the six downstream datasets. The gate is not a uniform mixer. The first layer leans on the temporal path (α = 0.745), layers 4–7 lean on the spatial path (β = 0.686–0.707), and late layers return toward the temporal path (α = 0.651 at layer 19). So the gate learns a different temporal/spatial balance at each depth. The rest of this appendix asks whether the model actually depends on these values.
Figure 4: AXON axis gate across the 22 encoder layers: gate weights averaged over all six downstream datasets (mean ± σ), on validation batches with the frozen encoder. Temporal-heavy in the first layer, spatial-heavy in the middle layers, and back toward temporal in the late layers.
16
H.2
Gate intervention diagnostics
We override the axis gate while keeping the pretrained encoder frozen, then run linear probe evaluation (full dataset splits). Each intervention replaces the learned per-token gate [αi , βi ] with a modified version that removes a specific component of the routing, isolating what the gate actually contributes. We group the seven interventions by the question they answer. Is axis factorization itself necessary? • temporal only: Force [αi , βi ] = [1, 0] for all tokens. Only the temporal path contributes. • spatial only: Force [αi , βi ] = [0, 1]. Only the spatial path contributes. Does soft mixing matter, or can the model commit to one axis? • hard argmax: Take the learned gate, find the dominant axis, and set it to 1.0 with the other at 0.0. E.g., [0.6, 0.4] → [1.0, 0.0]. • uniform: Force [αi , βi ] = [0.5, 0.5] for all tokens. No routing at all equal weight to both paths. Is the gate’s value per-token content routing, or a depth/position schedule? • layer mean: Calibrate over 50 forward passes to compute the average gate per layer, averaging across all tokens and batches. At inference, every token in layer ℓ receives the layer-ℓ mean gate, regardless of content. Tests whether the depth schedule alone is sufficient. • position mean: Calibrate the average gate per (layer, channel, time-patch) position. At inference, a token at position (c, t) in layer ℓ receives the calibrated mean for that position, regardless of signal content. Tests whether position-aware routing adds value beyond the depth schedule. • shuffled: Run the gate MLP normally to produce [αi , βi ] for each token, then randomly permute the gate values within each layer. The marginal distribution of gate values per layer is exactly preserved, but the token↔gate correspondence is destroyed. Tests whether it matters which token gets which gate value. Table 14 reports summary statistics; Figure 5 shows per-dataset results. Table 14: Gate intervention diagnostics under frozen-encoder LP. Mean is over all 6 downstream datasets. Scores are balanced accuracy on the test subjects at the validation-selected epoch, averaged over 3 downstream seeds for five datasets; siena uses a single seed and the class-balanced training loader. Gate mode
Mean BAC
∆ vs learned
0.571 0.539 0.559 0.567 0.564 0.501 0.471 0.446
— −5.6% −2.1% −0.8% −1.3% −12.2% −17.5% −22.0%
Learned (reference) Uniform (1/2, 1/2) Layer mean Position mean Shuffled Hard argmax Temporal only Spatial only
Protocol note. The learned-gate reference is 0.571 in this intervention setting, compared with the main AXON LP score of 0.579 in Table 1. All intervention results are measured relative to this within-table learned reference.
H.3
Token diversity and content dependence
We compute two scalar metrics to characterise within-layer gate variation. Dtoken measures how different individual token gates are from the layer mean; Dcontent subtracts out fixed channel/time position effects, isolating variation that depends on the signal content of the current input. Dcontent peaks sharply at layers 9 and 12 (Figure 6), indicating that signal-driven routing is concentrated at mid-depth rather than distributed uniformly across the encoder. However, the gate intervention results (Table 14) show that this token-level variation is not the dominant source of downstream gain: shuffled gates drop only −1.3% vs. learned. 17
Figure 5: Per-dataset gate intervention heatmap (frozen-encoder LP, 6 datasets; same protocol as Table 14). Cell colour is the change from the learned gate. The learned gate outperforms uniform and hard-argmax routing on most datasets, but layer-mean, position-mean, and shuffled gates stay close to learned, indicating that exact token-gate alignment is not the dominant effect.
Figure 6: Dtoken and Dcontent across 22 encoder layers. Peaks at layers 9 and 12 show that some layers exhibit genuine token-level gate diversity. However, the intervention results above show that this variation is not the dominant source of downstream gain: shuffled gates (which preserve the gate distribution but destroy token-gate alignment) drop only −1.3% vs learned.
H.4
Analysis of the AXON-TokenGated vs AXON-Fixed gap
Giving every token in a layer that layer’s mean gate costs only −2.1% (Section H.2). So at test time, one mixing weight per layer is almost enough. But AXON-Fixed learns exactly one weight per layer, and it reaches only 0.558 Mean LP, against 0.578 for AXON-TokenGated. Why? One guess is initialisation: AXON-Fixed starts from a neutral mixture and may never find the right weight for each layer. We tested this. We trained AXON-Fixed again, this time starting each layer’s weight at the value the AXON gate had learned for that layer (Section H.1). Nothing else changed. Table 15 shows the result: 0.555 Mean LP, the same as before (0.558), still far below 0.578. The right starting point does not help. So initialisation is not the reason. What is left is how the two models train. With a per-token gate, tokens in the same layer get different mixtures during training, so the temporal and spatial paths are trained on more varied signals. At the end of training the exact per-token values no longer matter much (shuffling them costs only −1.3%), 18
but the paths they trained are better. Dropout works the same way: it does nothing at test time, but it changes the weights that training ends with. Table 15: AXON-Fixed re-trained with each layer’s weight initialised to AXON’s learned per-layer balance. For comparison: AXON-Fixed with neutral initialisation reaches 0.558 Mean LP, AXONTokenGated 0.578.
H.5
Dataset
LP BAC
FT BAC
motor workload hmc siena adftd bcic
0.354 0.585 0.673 0.878 0.557 0.283
0.598 0.714 0.732 0.873 0.588 0.311
Mean
0.555
0.636
Temporal scale gate routing
Figure 7 shows the per-layer routing between the short-window (Ws = 5) and full-window (Wl = −1) temporal branches across all 22 encoder layers. The scale gate strongly favours the long-window branch (≈79% routing mass), consistent with the initialisation bias toward Wl . A small subset of layers routes appreciable mass to the short window, suggesting that local transient structure is selectively useful at those depths.
Figure 7: Scale gate routing: short window (W = 5) vs. full window (W = ∞) across 22 layers. The model usually favours broad temporal context, but uses short temporal windows in a few selected layers. The two-scale temporal design contributes +0.001 Mean LP and +0.013 Mean FT over the single-scale variant; the scale gate is a secondary improvement.
I
Interpreting the Attention
I.1
Does the trained dense model use the off-axis pairs that AXON removes?
For a query token in the modal C = 21, T = 11 grid, 200 of the 231 tokens share neither its channel nor its time step. We call these off-axis pairs. An untrained dense model spreads its attention uniformly, so about 87% of its attention starts on off-axis pairs. The question is how much of that the dense model learns to remove during pretraining. We measured it on the motor imagery grid (C = 64, T = 4), because motor imagery is the task with AXON’s largest gain and the 64-channel layout makes the pair types easy to separate; there the uniform off-axis share is 73.8%. After pretraining, attention within the same time step rises to 62.4% and the off-axis share falls to 30.9% on average 19
(Table 16). But it never goes away: the first layer still places 63% of its attention off-axis, almost the untrained value, and the last layer drifts back to 46% (Figure 8). The dense model learns the axis structure only partly and unevenly across depth. AXON assigns zero off-axis attention within a layer by construction. Table 16: Relation-type attention mass (%) in the dense baseline on the default motor mv img downstream grid (C = 64, T = 4), averaged across all 22 layers. After MAE pretraining, spatial mass rises and off-axis mass drops, but residual off-axis routing remains. AXON assigns zero one-layer off-axis mass by construction. Condition
Self
Temporal
Spatial
Off-axis
Uniform reference Dense (trained) AXON (by construction)
0.39 4.65 —
1.17 2.06 100 (temporal path)
24.6 62.4 100 (spatial path)
73.8 30.9 0
Figure 8: Relation-type attention mass per layer in the dense baseline (C = 64, T = 4). Top: Random initialisation matches the uniform complete-graph prior (dashed lines). Bottom: After MAE pretraining, dense attention discovers spatial dominance in mid-depth layers but fails to fully suppress off-axis interactions. Early and late layers retain up to 63% off-axis mass, leaving off-axis interactions that AXON removes by construction. I.2
Does AXON rely on the electrodes that physiology predicts?
The task is four-class motor imagery (motor mv img, Table 8). In each trial the subject imagines moving the left hand, the right hand, both fists, or the feet. The body is controlled from the opposite side of the brain: the left hand from the right hemisphere, the right hand from the left. Both fists use both sides. The feet are controlled from the midline. So for each class we know which electrodes a good classifier should be using. We take the trained AXON motor-imagery classifier and its held-out test subjects. We cover the seven electrodes over the left motor cortex, so the model gets no signal from them, and measure how much each class’s recall changes. Recall is the fraction of a class’s trials that the model labels correctly. Then we do the same for the seven electrodes over the right motor cortex. We chose the electrodes and the measure before looking at any result. 20
If the classifier uses the correct electrodes, covering one side should mainly hurt the opposite hand, and should not hurt both fists or feet in a one-sided way. If it used all electrodes alike, covering either side would hurt all four classes alike. Table 17 shows the first pattern. Covering the left side drops right-hand recall by 0.101 and does not hurt the left hand (+0.058). Covering the right side drops left-hand recall by 0.127 and does not hurt the right hand (−0.004). Both fists and feet show no one-sided change. So the drop is not just “fewer electrodes, worse accuracy”: each hand’s decision depends on the electrodes over the opposite hemisphere, exactly where physiology says it should. We use this test rather than an attention map because it changes the input and watches the decision; an attention map only shows where the weights point. Table 17: Change in recall after covering the left or right motor strip (motor imagery, held-out subjects); negative means worse. Imagined movement
Cover LEFT strip
Cover RIGHT strip
+0.058 −0.101 −0.027 +0.013
−0.127 −0.004 −0.001 +0.052
Left hand Right hand Both fists Feet
I.3
Does the accuracy depend on particular electrodes?
Deleting randomly chosen electrodes at test time, AXON stays ahead of dense (balanced accuracy) at every level, from 0.377 vs. 0.321 intact to 0.274 vs. 0.263 with 94% of electrodes removed. Here the classification head is trained on mean-pooled frozen features rather than through our full evaluation harness, so absolutes differ from Table 1 and only the model-to-model comparison is meaningful. This matters clinically, where reduced montages and failed electrodes are routine.
J
Cross-Modal Generalization (AudioMAE)
Gap to published AudioMAE. The absolute mAP values are not directly comparable to those reported by Huang et al. [2022]. For example, our best full AudioSet-2M FT result is 14.29 mAP, whereas published AudioMAE-style results report much higher absolute mAP under a substantially different training recipe (e.g., 37.0 on AudioSet-20K). This gap reflects five controlled differences: (1) spectrogram resolution (128 vs. 8 frequency bins, a 16× reduction that brings the frequency axis to a size comparable to the EEG electrode axis); (2) pretraining compute (4 epochs vs. 32 epochs, ∼9× fewer sample-views); (3) model capacity (d = 512 vs. d = 768, ∼50% fewer parameters); (4) decoder depth (4 vs. 16 layers); and (5) FT augmentation (no Mixup, SpecAugment, or DropPath). Crucially, the factorized and dense variants share all five of these constraints, so the within-setup comparison is valid. J.1
Small-scale preliminary results (18K AudioSet clips)
Before scaling to 200K clips, we verified factorized attention at 1% of AudioSet (∼18K clips). The same four-variant design (Dense, Factorized fixed gate, Factorized token gate, Token+Global) was trained for 33 epochs with identical optimizer and schedule. Table 18: Audio results at 1% scale (18K AudioSet clips) Model Dense Factorized (fixed gate) Factorized (token gate) Token + Global
AudioSet FT mAP
AudioSet LP mAP
ESC-50 LP Acc
SC LP Acc
6.76±0.07 10.36±0.12 10.48±0.33 10.09±0.14
1.18±0.01 1.21±0.01 1.32±0.01 1.40±0.01
20.50±0.43 21.33±2.75 20.67±0.63 21.50±0.20
11.23±0.76 11.70±0.46 12.26±0.03 12.64±0.10
All factorized variants improved AudioSet FT mAP by +49–55% relative over the dense baseline even at this small scale, demonstrating that the factorized advantage is present from the smallest dataset scale tested. 21
J.2
200K-scale results
Table 19: Audio results at 200K scale (∼10% of AudioSet-2M) Model
AudioSet FT mAP
AudioSet LP mAP
ESC-50 LP Acc
SC LP Acc
11.04±0.25 13.52±0.15 13.75±0.20 13.99±0.15
3.61±0.00 3.76±0.03 4.09±0.04 4.40±0.04
40.33±0.51 43.42±0.94 44.42±0.42 44.75±1.27
17.71±0.12 19.88±0.08 21.08±0.10 20.00±0.18
Dense Factorized (fixed gate) Factorized (token gate) Token + Global
At 200K, the factorized advantage persists across all four metrics. Token+Global wins 3 of 4 metrics but underperforms the token-gate model on SpeechCommands LP (20.00 vs. 21.08). The SpeechCommands ordering varies across scale: Token+Global is ahead at 18K (Table 18: 12.64 vs. 12.26), behind at 200K, and approximately tied with the token-gate model at 2M (Table 2: 21.87 vs. 21.96). We therefore avoid drawing a stable architectural conclusion from this single metric and focus on the consistent factorized-vs-dense improvement.
K
Cross-Modal Gate Mechanism Analysis
To test whether the gate intervention findings from EEG (§H.2) are EEG-specific or reflect a general property of axis-factorized attention, we run the same seven gate overrides on the audio AXONTokenGated model (200K AudioSet pretraining). The encoder is frozen; only the linear probe head is trained. We evaluate on ESC-50 and SpeechCommands v2. Table 20: Audio gate intervention results. The same seven overrides from the EEG analysis are applied to the frozen audio AXON-TokenGated encoder. ∆ is the absolute change in accuracy relative to the learned baseline. ESC-50 Acc
SC Acc
ESC-50 ∆
SC ∆
Learned (baseline)
44.50
20.93
—
—
Shuffled Layer-mean Uniform (1/2, 1/2) Position-mean Hard argmax Frequency-only (α = 0) Temporal-only (β = 0)
44.50 44.25 44.00 43.50 39.00 29.75 23.75
20.98 20.71 21.01 20.59 20.39 18.03 13.64
+0.00 −0.25 −0.50 −1.00 −5.50 −14.75 −20.75
+0.05 −0.22 +0.08 −0.34 −0.54 −2.90 −7.29
Intervention
Cross-modal comparison. Table 21 compares the mean relative degradation of each intervention across EEG (6 tasks, balanced accuracy) and audio (2 tasks, accuracy). Both columns report the mean relative change (∆/baseline) × 100. The interventions show a similar broad pattern across mannas.ais: removing either axis or replacing soft mixing with hard selection is more damaging than averaging or shuffling the gates. The exact ordering and magnitudes differ: spatial-only is the worst override on EEG but temporal-only is the worst on audio, and uniform mixing costs 5.6% on EEG but almost nothing on audio. Audio layer-mean gate profile. The calibrated per-layer gate means reveal an interpretable depth schedule. Early layers favour the frequency axis (α1 = 0.422, frequency-heavy), mid layers are approximately balanced (α3–5 ≈ 0.485), and late layers shift toward the temporal axis (α10 = 0.564, α11 = 0.556). This is the opposite direction from EEG, where early layers are temporal-heavy (α0 = 0.745) and mid layers are spatial-heavy. The reversal is consistent with mannas.ai structure: early audio layers capture spectral features (pitch, harmonics) that require cross-frequency integration, while late layers capture temporal dynamics (onsets, rhythm) that require cross-time integration. In EEG, the early temporal bias captures fast transient features (spikes, ERD onset) before spatial mixing integrates across electrodes. 22
Figure 9: Audio gate intervention heatmap (∆% vs. learned baseline). The pattern mirrors EEG: single-axis routing is catastrophic, while shuffled and layer-mean gates are indistinguishable from the learned gate. Table 21: Gate intervention mean relative degradation (%): EEG vs. audio. Both columns report the mean relative change from the learned baseline. The broad pattern is shared across mannas.ais; the exact ordering and magnitudes differ. Intervention Shuffled Layer-mean Position-mean Uniform Hard argmax Spatial/Freq-only Temporal-only
EEG Mean ∆%
Audio Mean ∆%
−1.3 −2.1 −0.8 −5.6 −12.2 −22.0 −17.5
+0.1 −0.8 −1.9 −0.4 −7.5 −23.5 −40.7
Uniform robustness gap. The most notable cross-modal difference is the uniform intervention: −5.6% in EEG but only −0.4% in audio. This indicates that the audio schedule is flatter, the per-layer gate values range from α = 0.422 to 0.564 (range 0.14), compared to α = 0.294 to 0.745 (range 0.45) in EEG. Audio representations benefit nearly equally from both axes at all depths, while EEG requires stronger layer-varying axis preferences. This is consistent with audio spectrograms containing genuine cross-axis harmonic structure at all levels, whereas EEG temporal and spatial dynamics are more separable. Summary. The gate interventions show the same broad pattern in both mannas.ais: both axes and soft mixing are necessary, while the dominant useful gate structure is the learned per-layer temporal/spatial balance. Token-level content routing is secondary rather than dominant in both EEG 23
(−1.3% shuffled) and audio (+0.1% shuffled). The token gate serves as a training mechanism that discovers an appropriate layer-wise axis schedule, and the optimal schedule direction differs between mannas.ais (temporal-first in EEG, frequency-first in audio). K.1
Limitations
Several limitations should be noted. First, the theoretical support for removing the global path rests on empirical ablations (Table 11) rather than a formal information-theoretic proof for nonlinear masked autoencoders; such a treatment remains open. Second, the audio experiment uses a reduced spectrogram resolution (8 frequency bins vs. AudioMAE’s 128), limiting the absolute performance achievable; while the within-setup comparison is valid, the factorized advantage under full spectral resolution has not been verified. Third, the pretraining corpus pools clinical EEG recordings from a limited number of sources; performance on substantially different populations or recording protocols has not been evaluated. K.2
Future Work
Priority directions include: (1) repeating the AudioMAE experiment at full spectrogram resolution (128 frequency bins) to determine whether the factorized advantage persists when spectral detail is not bottlenecked; (2) investigating whether a task-conditioned gate temperature could resolve the temporal locality tradeoff across clinical and BCI tasks simultaneously; and (3) evaluating AXON on additional downstream mannas.ais such as sleep staging with polysomnography and intracranial EEG.
24