Conceptio › Archive › arXiv CS
arXiv CSopen access

Relational Attention for Data-Efficient Language Modeling

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Relational Attention for Data-Efficient Language Modeling Adrian Brasoveanu UC Santa Cruz, [email protected]

Ece Takmaz Utrecht University [email protected]

arXiv:2609.20530v1 [cs.CL] 17 Sep 2026

Abstract

We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level (“sensory”) lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM’s data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT’s three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPEbased, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100Mword) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard’s NLPtask subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.

1

Jakub Dotlačil Utrecht University [email protected]

Introduction

Standard Transformer language models (Vaswani et al., 2017; Radford et al., 2019) face two structural limitations in the low-data BabyLM regime. First, ordinary self-attention entangles relational information with high-dimensional object-level features: the query–key dot product is derived from the inputs / sources / objects, and can be understood as encoding the relation between the objects and controlling which source / object is selected (via the softmax) based on that relational information, but the routed value is a feature vector derived from that same source / object. This entanglement is at odds with the division of labor in, for example, formal semantics between lexical semantics, which makes word-specific, object-like contributions to sentential meaning (truth conditions), and compositional semantics, which makes structure-based, relation-like contributions to meaning (Montague, 1973; Dowty et al., 1981; Partee, 1995). The ability to separate relational structures from sensory features is evident in infants as young as seven months, who can readily generalize algebraic rules (relational patterns: ABA vs. ABB) to novel vocabularies (objects) (Marcus et al., 1999). Infants generalize this form of relational (symbolic) reasoning quickly, from few ‘training’ examples, and apply it systematically to novel out-of-trainingsample (OOTS) objects. In contrast, self-attentionbased transformers require a lot more data, and still struggle with OOTS generalization (Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). A second issue with self-attention is that, unlike RNNs, it allows a model to look back at any past token, so the architecture lacks an inherent pressure to compress sequence history into a dense, unified summary that can support OOTS generalization. This contrasts with human language processing, which is subject to a “Now-or-Never” bottleneck (Christiansen and Chater, 2016): linguistic input

must be rapidly chunked and compressed into hierarchical representations before fleeting sensory memory decays. This forces eager incremental processing that drives expectations in rational parsing (Hale, 2011) and cue-based memory retrieval (Lewis and Vasishth, 2005). In this paper, we present Relational BabyLM, a system submission to the BabyLM 2026 challenge (Choshen et al., 2026) that combines two cognitively motivated inductive biases in a single decoder-only Transformer. The primary architectural contribution is (our own reimplementation of) the Dual Attention Transformer (DAT; Altabaa and Lafferty, 2025), which augments sensory selfattention with a parallel relational attention mechanism: source selection works as usual, but what is retrieved is an explicit pairwise relation vector together with an abstract symbol identifying the source (Section 2.1). DAT’s data-efficiency advantages have been largely established on purely relational tasks, and the evidence for its usefulness in language modeling is limited to modest perplexity gains (Altabaa and Lafferty, 2025). Whether relational attention helps a language model under a broad linguistic evaluation has remained an open question investigated by our present submission. The complementary training intervention is Next-Latent Prediction (NextLat; Teoh et al., 2026), which trains a small auxiliary dynamics model to predict the next hidden state from the current one and the actual next token (RNN-style), pushing the hidden states toward a belief state, a sufficient statistic of the past for predicting the future similar to dynamic semantics info states (Kamp, 1981; Heim, 1982; Groenendijk and Stokhof, 1991). The dynamics model is discarded after training, leaving the base architecture and autoregressive inference unmodified (Section 2.4). We submit models to the strict (100M-word) and strict-small (10M-word) tracks, and run a controlled experimental matrix on the strict-small track to separate the two inductive biases. The evaluation reveals an asymmetry between them. 1. Architecture is the dominant factor for structural generalization (Sections 4.2 and 4.6). DAT configurations outperform matched standard Transformers on BLiMP, with the largest gains on structural domains such as island effects and subject–verb agreement, though a head-ratio sweep shows that enlarging the relational stream, rather than

merely having one, does not help further. 2. The training objective is a secondary but significant factor for cognitive alignment (Section 4.3). NextLat models explain more variance in human reading measures than nexttoken-prediction models, and improve 5 of 7 fine-tuned (Super)GLUE task accuracies. 3. The symbol-retrieval mechanism matters (Section 4.4). In a fully powered fiveseed comparison across seven mechanisms, relational-symbolic variants underperform and the rest are indistinguishable, so our new RoPE-based relative symbols match a learned relative symbol library at no parameter cost. On the strict (100M-word) track, our two strongest models rank among the top ten entries on the official leaderboard at the time of writing, with our best model 6th of 55 overall and 3rd of 55 on the NLP-task subset, and both outperform the official GPT-2 baseline on most benchmarks; on the strict-small (10M-word) track, our NextLat submission ranks 8th of 127 on BLiMP and 6th on EWoK (Section 4.1).

2

Model

The Dual Attention Transformer (DAT) (Altabaa and Lafferty, 2025) separates two kinds of information that are entangled in ordinary self-attention. Standard self-attention routes sensory information: it selects source objects by comparing queries and keys, then routes their feature vectors. DAT adds relational attention, whose routed values explicitly encode relations between a receiving object and source objects. The resulting layer contains both sensory heads and relational heads, allowing the model to retrieve object-level features and relationlevel information in parallel. 2.1

Sensory and Relational Attention

Let x = (x1 , . . . , xn ) ∈ Rn×d be a sequence of object representations, where n is the sequence length and d is the model (embedding) dimension, so that each xi ∈ Rd . Standard sensory attention from receiver xi to context x is P Attn(xi , x) = nj=1 αij xj Wv , (1) n αi = Softmax ⟨xi Wq , xj Wk ⟩ j=1 , (2) where Wq , Wk , Wv are learned query, key, and value projection matrices and αij is the j-th component of αi . The attention weights αij encode a

selection criterion, but the retrieved values xj Wv remain sensory representations of source objects. Relational attention keeps the selection mechanism but changes what is retrieved. Instead of routing only xj Wv , the receiver retrieves a relation vector between xi and xj , tagged by a symbol sij identifying the source: RelAttn(xi , x) = n X  αij r(xi , xj ) Wr + sij Ws ,

(3)

j=1

where the pairwise relation vector r(xi , xj ) ∈ Rdr stacks dr learned comparison channels,  dr  r(xi , xj ) = σrel ⟨xi Wqrel,ℓ , xj Wkrel,ℓ ⟩ ℓ=1 , (4) each channel ℓ using its own relational projections Wqrel,ℓ , Wkrel,ℓ , and σrel being a relation activation function (e.g., identity, sigmoid, tanh, or softmax). The symbol sij ∈ Rds identifies the source as an abstract symbol (Section 2.2). The projections Wr ∈ Rdr ×dv and Ws ∈ Rds ×dv map the relation and symbol into a common output space Rdv , so their contributions can be summed. Thus, the “message” from source j to receiver i contains a source-receiver relation r(xi , xj ) and a symbol sij tagging the sender. In principle, Transformers with only selfattention or only relational attention could be used for language modeling. The Dual Attention Transformer combines the two. For nsa h sensory heads ra and nh relational heads, the combined output for position i concatenates the two groups and applies an output projection. The SA/RA head split controls the allocation to sensory versus relational retrieval: a 6SA/6RA split gives half to each stream, while 9SA/3RA gives three quarters to sensory attention. We instantiate DAT with pre-normalized causal decoder blocks, applying the same decoder mask to sensory and relational heads before the softmax over source positions. 2.2

Symbol Assignment

Relational attention tags each source j with a symbol that is receiver-conditioned, written sij . Some of the symbol “libraries” and associated retrieval mechanisms can depend on the source alone (sij = sj ), as with a symbol tied to the absolute position j, whereas other symbol libraries make it depend on both, for example, relative position symbols for the receiver–source offset j − i. The original DAT

design (Altabaa and Lafferty, 2025) offers three mechanisms—symbols indexed by absolute position, symbols indexed by the receiver–source offset, and a learned symbol library retrieved via symbolic attention (or a relational-symbolic variant)—and uses symbolic attention for its language models. We add a fourth, RoPE-based relative symbols: in the original work (Su et al., 2023) rotary encodings act on the queries and keys as a positional encoding, whereas here they generate the relative symbol (the “value”) itself. The symbol mechanisms we study in a controlled five-seed comparison are: • Learned relative symbols (relative): a learned symbol library indexed by the relative offset j − i, clipped to a maximum distance ∆. Each source is tagged by a receiverconditioned relative symbol. relative symbols • RoPE-based (relative_rope, new): instead of a learned offset-indexed library, the relative symbol at offset j − i is generated by applying rotary position encoding (RoPE) to the clipped relative distance. This gives a deterministic relative symbol that does not require learning a separate symbol library. • Learned positional symbols (positional): a learned symbol library indexed by absolute source position j, so the symbol sj depends only on the source position in the sequence. positional symbols • Sinusoidal (positional_sinusoidal): the same as learned positional symbols, but the position-indexed symbol library is replaced by a fixed sinusoidal positional encoding, as in one of the Abstractor experiments (Altabaa et al., 2024). • Symbolic attention (symbolic): maps each object to a convex combination of abstract symbols from a learned library S = (s1 , . . . , sns ) via learned feature templates. • Relational-symbolic (relsymbolic(_n4)): first computes a local relation profile for each position over a causal neighborhood of size k (default 2, or 4), then applies symbolic attention to that profile. 2.3

Rel. Attn. Type: RA, DisRCA, and RCA

We compare three relational attention types, which differ in (i) whether the attention weight itself serves as the relation, and (ii) what is routed. RCA

(Altabaa et al., 2024) is self-attention P where the values are replaced by symbols: x′i = j αij sij , with αij = σrel (⟨xi Wq , xj Wk ⟩). The attention weight itself is the relation, and the routed values are symbols rather than sensory features. The Abstractor treats σrel as a configurable hyperparameter, softmax in most of its experiments but element-wise activations in others, since softmax can mask relevant information. Our runs use identity (architecture comparison, head-ratio sweep, submitted models) or sigmoid (symbol-retrieval comparison; Appendix B). DisRCA (Altabaa and Lafferty, 2025) is an intermediate design between RCA and RA. A softmax attention weight αij selects sources, which are gated by a separately computed per-head relational compatibility rij —the construction of Eq. 3 taken as one scalar channel per relationalP head rather than as a dr -dimensional vector: x′i = j αij rij sij . The elementwise product creates an AND-like condition—a source is routed only if it is both selected and relationally compatible—and only the symbol is routed. Full RA (Altabaa and Lafferty 2025; Section 2.1 above) separates selection from relation. Two independent sets of query/key maps compute (i) a softmax attention weight αij for selecting sources and (ii) a separate relation P vector r(xi , xj ). The output routes both: x′i = j αij r(xi , xj ) Wr + sij Ws . 2.4

Next-Latent Prediction and Belief States

Self-attention’s unrestricted look-back gives a Transformer no inherent pressure to compress sequence history into a dense, unified summary like RNNs do. To add this pressure with the goal of OOTS generalization, we augment the training objective with Next-Latent Prediction (NextLat; Teoh et al., 2026). Let ht be the final-layer hidden state at position t and pθ (· | ht ) the LM head that maps it to a distribution over the vocabulary. A small latent dynamics model predicts the next hidden state from the current one and the next token, ĥt+1 = ht + δψ ([xt+1 ; ht ]), where the update δψ is a LayerNorm over [xt+1 ; ht ] followed by a three-layer GELU MLP of width 2d; the skip in the equation is the only residual connection. Dropout (0.1, as elsewhere in the model) is applied to ht entering δψ , but not to the skip path. Alongside the usual next-token cross-entropy LNTP , two auxiliary losses compare that prediction against the true ht+1 : a Smooth L1 loss Lh on the hidden state itself, and a token-space divergence  sg LKL = DKL psg (· | sg[h ]) ∥ p (· | ĥ ) . t+1 t+1 θ θ

Targets are detached (sg[·]) to prevent representational collapse, and the LM head is frozen inside the KL term, so the only remaining gradient path runs through the predicted state ĥt+1 , shaping the backbone’s representations rather than the LM head. The total objective is L = LNTP + λh Lh + λKL LKL ,

(5)

with λh = 1.0 and λKL = 0.5 throughout; an auxiliary cross-entropy on the predicted state’s emissions is available but carries weight zero in every reported run. We use the one-step setting, which is what the belief-state result requires; the original NextLat paper additionally studies multi-step rollouts in service of speculative decoding. Unlike the token-level losses, Lh applies at every position whose prediction does not cross a document boundary, so belief-state structure is shaped even while processing context. These losses encourage the hidden states to converge toward a belief state, a sufficient statistic of the past for predicting the future. The auxiliary dynamics model is discarded after training, preserving the unmodified base architecture and standard autoregressive inference.

3

Pretraining and Evaluation

3.1

Challenge Setting and Corpora

We train on the BabyLM strict-small and strict tracks, which provide 10 million words and 100 million words respectively. Both tracks consist of developmentally plausible English text drawn from six sources: CHILDES, OpenSubtitles, Simple Wikipedia, Gutenberg, BNC spoken, and Switchboard (Choshen et al., 2026; Warstadt et al., 2023). The strict-small regime is the most data-limited BabyLM setting and is therefore the most diagnostic of inductive-bias effects. We tokenize with byte-pair-encoding (BPE) tokenizers in the GPT-BERT recipe (Charpentier and Samuel, 2024), each with a vocabulary of 16,384. The strict-small track and the 18-layer strict model use the tokenizers distributed with the official BabyLM GPT-BERT baselines (the 10Mword and 100M-word releases, respectively); the 16-layer wide (1024 hidden dim) strict model uses a 2026 retraining of the 100M-word tokenizer on the 2026 strict corpus. Training sequences are packed to a length of 512 tokens (264 for the symbolretrieval comparison; Appendix B). To preserve the short-utterance structure of child-directed input, we segment sequences at end-of-sequence (EOS)

boundaries: tokens after an EOS boundary cannot attend to tokens before it, and the auxiliary dynamics model is never penalized for predicting latent states across EOS: causal next-token prediction (NTP) training with utterance-level segmentation. 3.2

Model Configurations and Pretraining

All strict-small models share a 12-layer, 768dimensional decoder-only backbone with a BPE vocabulary of 16,384, rotary positional encodings, pre-LayerNorm blocks, and GELU or SwiGLU feed-forward activations depending on configuration. The standard Transformer baseline uses 12 self-attention heads; DAT replaces these with nsa h sensory and nra h relational heads, with the split controlling the sensory/relational allocation. LM-head tying varies independently of the objective: NextLat runs use an untied head, matching the source NextLat setup, while NTP runs appear with both tied and untied heads. Tied models have ≈123.3M parameters and untied models ≈135.9M, the untied head adding ≈12.6M parameters (the auxiliary NextLat dynamics model adds ≈5.9M parameters and is discarded after training). All models train for 10 epochs with gradient clipping at 1.0. Batch sizes, learning rates, and head splits vary across the three analysis families and are summarized in Appendix B. The strict-track submissions use the Muon/LambW recipe of Section 3.3, and their full specifications are in Appendix C. 3.3

Optimization

Our experiments use two optimization recipes. The strict-small analysis families use AdamW (Loshchilov and Hutter, 2019) with no weight decay, β1 =0.9, β2 =0.999, and the family-specific learning rates of Appendix B. The strict-track submissions use Muon (Jordan et al., 2024; Liu et al., 2025) for matrix-valued hidden-layer parameters and LambW (our own decoupled-weight-decay variant of LAMB (You et al., 2020), which relates to LAMB as AdamW relates to Adam) for embeddings, the LM head, symbol tables, normalization parameters, and biases. Muon applies a Newton– Schulz approximation to orthogonalize momentum updates; we use a Muon learning rate of 0.02, momentum of 0.95, and 5 Newton–Schulz steps. LambW retains LAMB’s layerwise trust ratio while applying weight decay independently of the adaptive update, following the decoupling principle of AdamW (Kingma and Ba, 2015; Loshchilov and Hutter, 2019). We use a LambW learning rate of

10−3 and weight decay of 0.01. We compare the two recipes head-to-head in one controlled stricttrack experiment, with AdamW at a learning rate of 5 × 10−4 (Table 7, Appendix E); because it jointly changes optimizer, learning rate, and weight decay, the comparison reflects the whole recipe rather than the optimizer alone. 3.4

Experimental Comparisons

Our experimental matrix comprises 140 training/evaluation rows across 47 experiment groups, organized into three analysis families: (i) an architecture comparison (standard Transformer vs. DAT) with 77 rows, (ii) a complete RCA/SwiGLU symbol-retrieval comparison at 264-token context window with 35 rows (7 symbol conditions × 5 seeds), and (iii) an RCA SA/RA head-ratio sweep with 28 rows (7 ratios at 512-token context with 3 seeds, plus one 264-token run per ratio). All models use the strict-small (10M-word) track. SwiGLU was chosen early in an unreported comparison with GELU, and was held fixed for all the submitted models, as well as the symbol-retrieval and headratio comparison families; the architecture comparison includes both GELU and SwiGLU runs, with activation partly coupled with symbol mechanism. The architecture comparison varies four factors in a partially crossed design: Base architecture— standard self-attention with 12 heads vs. DAT with 9 sensory + 3 relational, or 6 + 6 heads; Training objective—next-token prediction (NTP) vs. NextLatent Prediction (NextLat); LM head—tied vs. untied; and Relational attention type—full relational attention (RA), relational cross-attention (RCA), or disentangled RCA (DisRCA). 3.5

Evaluation

We evaluate on the BabyLM zero-shot benchmark suite, which probes distinct linguistic and cognitive dimensions without task-specific fine-tuning: grammatical phenomena (BLiMP, BLiMP supplement, Warstadt et al. 2020), conceptual property knowledge (COMPS, Misra et al. 2023), discourse tracking and world knowledge (Entity Tracking, Kim and Schuster 2023, EWoK, Ivanova et al. 2025, GlobalPIQA, Chang et al. 2025), and humanlikeness measures (fit to human reading-time and EEG data, and age of acquisition). We also evaluate selected models on the (Super)GLUE fine-tuning suite. Detailed benchmark definitions are provided in Appendix D.

Our claims rest on two statistical models, each fit on the strict-small (10M-word) track. Binomial GLMMs on BLiMP. We fit binomial generalized linear mixed-effects models (GLMMs) on BLiMP pseudo-items (successes out of items per pseudo-item). A baseline-vs-DAT model compares the standard Transformer against the 9SA/3RA RCA/SwiGLU DAT configuration, with fixed effects for base architecture, LM-head tying, and training objective; an internal-DAT model adds fixed effects for relational attention type, head split, feed-forward activation and symbol-position setting, and context length. All models use Type III Wald χ2 tests with sum-to-zero contrast coding, and crossed random intercepts for experiment (which identifies the seed and configuration), BLiMP subtest, and pseudo-item ID. Post-hoc contrasts use Tukey-adjusted pairwise comparisons. Throughout, the architecture comparison’s factors are only partially crossed (tied/NextLat is unobserved), so contrasts involving a missing cell are model-based extrapolations. Reading regression. A cross-experiment linear mixed-effects model regresses the incremental R2 (∆R2 ) gained by adding model surprisal to baseline reading-time predictors, with crossed random intercepts for experiment and reading measure, and fixed effects for training objective, model type, LMhead tying, context length, and reading condition. The last of these is an evaluation-condition control: each experiment contributes two ∆R2 rows per measure, one from the base reading regression and one from a spillover regression that adds previousword predictors.

4

Experiments

4.1

BabyLM Submission

Table 1 summarizes our strict-track models and the official GPT-2 baseline under the official leaderboard evaluation. Our two strongest entries are DAT models trained with NextLat, the Muon/LambW recipe, and RoPE-based relative symbols, each on a two-phase curriculum that starts with line (utterance) EOS segmentation and switches to document EOS boundaries: an 18-layer model (768-dim, 9SA/3RA, RCA) trained with a curriculum of 6 line-EOS / 4 document-EOS epochs, and a 16-layer wide model (1,024-dim, 12SA/4RA, RA) trained with a curriculum of 2 line-EOS / 8 document-EOS epochs. Both submit-

ted checkpoints are exponential moving averages (EMA) of the training weights; Appendix C gives their full specifications. At the time of writing our two strongest entries rank 6th and 7th of 55 overall; our best, the 16-layer wide (1024-dimensional) model, also ranks 3rd of 55 on the NLP-task subset, which comprises all tasks except age of acquisition and fit to reading data. The wide model beats the baseline on eight of the nine benchmarks (all but GlobalPIQA) and holds the highest EWoK score on the strict-track leaderboard at the time of writing (59.54); the 18-layer model beats the baseline on seven of nine, and shows no reliable age-ofacquisition correlation (0.00; non-significant correlations are scored as zero, Appendix D) whereas the baseline’s correlation is significantly negative (−11.58). Its curriculum trades a little BLiMP accuracy for a large gain on the BLiMP supplement (+8.5 points over the non-curriculum 18L model). Note that these are not controlled parameter or training-recipe matched comparisons: the GPT-2 baseline has 98M parameters trained with AdamW, while our entries range from 123M to 304M parameters trained with Muon/LambW. A parameter-matched architecture comparison is provided in Section 4.2 below with models trained on the strict-small track. On the strict-small (10M-word) track, we submit two 12-layer DAT models with the same 9SA/3RA RCA architecture, SwiGLU activations, RoPEbased relative symbols, and Muon/LambW recipe: a base NTP model with a tied head (123M parameters) and a NextLat model with an untied head (136M parameters), both trained for 10 epochs at 512-token context with seed 1. At the time of writing they rank 60th and 53rd of 127 entries by Overall Average (37.91 and 38.16, versus 37.38 for the GPT-2 baseline), but place near the top on the structural and knowledge-oriented benchmarks: the NextLat model ranks 8th of 127 on BLiMP (72.34, versus 65.23 for the baseline) and 6th on EWoK (52.93). Relative to the base model, NextLat improves BLiMP (+1.9), EWoK (+2.7), reading (7.12 vs. 6.56), and (Super)GLUE (63.97 vs. 62.96); the two models differ in LM-head tying and therefore parameter count. Table 6 in Appendix E reports the full strict-small results. 4.2

Does Dual Attention Help?

For structural syntactic reasoning (BLiMP), architecture choice is the dominant factor. A binomial GLMM with fixed effects for base architecture

Model GPT-2 baseline (12L) DAT 12L (NTP) DAT 18L (NextLat) DAT 18L (NextLat, curric.) DAT 16L wide (NextLat, curric.)

BLiMP Supp. EWoK Ent. trk. COMPS PIQA GLUE∗ Reading

AoA Overall

NLP

74.73 65.00

54.37

16.91

55.85 36.62

67.75

6.93 −11.58

40.73 53.03

79.81 60.03

56.66

20.84

57.84 37.65

65.92

5.76 −12.02

41.39 54.11

80.62 62.01

56.99

20.93

58.36 34.72

68.90

6.49

0.00

43.23 54.65

78.94 70.51

57.25

20.47

58.69 32.72

69.95

6.22

0.00

43.86 55.50

79.49 70.96

59.54

20.89

59.22 36.17

71.43

7.10

−9.48

43.92 56.81

Table 1: Strict-track (100M-word) results from the official BabyLM leaderboard at the time of writing (55 entries). All DAT models use relative symbol retrieval with RoPE-based relative symbols (no learned symbol library; Section 2.2), SwiGLU activations, and the Muon/LambW recipe: the 12L model has 123M parameters (9SA/3RA, RCA, 768-dim, tied head, seed 0), the 18L models have 191M parameters (9SA/3RA, RCA, 768-dim, seed 1), and the 16L wide model has 304M parameters (12SA/4RA, RA with four relation channels, 1,024-dim, seed 1). Full specifications of the two top entries are in Table 4. PIQA is GlobalPIQA; GLUE∗ is the leaderboard’s (Super)GLUE fine-tuning average; Reading is the leaderboard’s composite human-likeness score, which aggregates eleven eye-tracking, self-paced-reading, and ERP measures (Appendix D); AoA is the age-of-acquisition score (Appendix D); Overall is the leaderboard Overall Average (equal weight across the nine benchmarks); NLP is the Overall Average restricted to the seven NLP-task benchmarks, excluding the two human-likeness measures (reading and age of acquisition, which form the leaderboard’s Human-like Average). The GPT-2 baseline is the official challenge baseline (98M parameters, AdamW). Bold marks the best value per column. The comparison with the GPT-2 baseline is not parameter or training-recipe matched.

(standard Transformer vs. DAT), LM-head tying, and training objective reveals a large and significant main effect of architecture (χ2 (1) = 139.43, p < 0.001; Table 11, Appendix E). Estimated contrasts show that DAT scores higher than the standard Transformer baseline across all four headtying × objective conditions, with advantages ranging from +0.083 to +0.288 on the log-odds scale (all p < 0.001; Table 2). Note that contrasts at factor combinations not present in the training matrix (such as tied/NextLat) are model-based extrapolations (Section 3.5). The best DAT configuration reaches 70.66% BLiMP against 68.16% for the best standard Transformer (Table 10, Appendix E).

Figure 1: BLiMP accuracy by linguistic category, comparing DAT and standard Transformer configurations. Gains are largest on structural domains (e.g., island effects, subject–verb agreement).

4.3 The advantage is not uniform across linguistic phenomena. A focused architecture × linguisticdomain interaction model shows a significant interaction (χ2 (12) = 573.07, p < 0.001): DAT yields larger gains on structural domains such as island effects, subject–verb agreement, and quantifiers than on lexical or semantic ones. Figure 1 shows BLiMP accuracy by linguistic category for DAT and standard Transformer configurations; Figure 2 in Appendix E gives mean BLiMP accuracy across configurations.

Does NextLat Complement Dual Attention?

While architecture governs structural accuracy, the training objective governs alignment with human cognitive processing. A cross-experiment reading regression models the incremental R2 (∆R2 ) gained by adding model surprisal to baseline reading-time predictors. The training objective is significant (t = 5.22, p < 0.001): NextLat models explain more variance in human reading measures than NTP models (Table 9, Appendix E). Model type (DAT vs. standard) is also significant (t = −2.61, p = 0.010), with standard Transform-

Contrast

LM head

Objective

Estimate

SE

z

p

baseline − DAT baseline − DAT baseline − DAT baseline − DAT

tied untied tied untied

NextLat NextLat NTP NTP

−0.288 −0.186 −0.185 −0.083

0.038 0.022 0.022 0.022

−7.49 −8.37 −8.33 −3.72

< 0.001 < 0.001 < 0.001 < 0.001

Table 2: Estimated contrasts (baseline − DAT) from the BLiMP GLMM, by LM-head tying and training objective. Negative estimates indicate that DAT scores higher than the standard Transformer baseline. All four contrasts are significant at p < 0.001 (Tukey-adjusted).

ers explaining slightly more reading-time variance than DAT models, and context length is significant (t = −7.08, p < 0.001). The matrix separates the objective from LM-head tying: its untied/NTP and untied/NextLat cells are matched on both, and NextLat gains 0.83 BLiMP points across them (Appendix E). Fine-tuned task performance. On the stricttrack (Super)GLUE benchmarks, NextLat improves 5 of 7 task accuracies, with the largest gains on MultiRC (+5.7 percentage points) and WSC (+3.8; Table 8, Appendix E). These two stricttrack models differ in LM-head tying and parameter count (≈ 12.6M). Architecture×objective interaction. The full baseline-vs-DAT model reveals significant interactions between base architecture and both head tying (χ2 (1) = 10.62, p = 0.001) and training objective (χ2 (1) = 10.81, p = 0.001): the size of the DAT advantage depends on the objective and on whether the LM head is tied, so the two interventions are not simply additive. 4.4

Which Relational Mechanisms Matter?

The RCA/SwiGLU 264-token symbol-retrieval comparison is the only fully-powered complete subset in our matrix: seven symbol conditions × five seeds (Table 15, Appendix E). A binomial GLMM on BLiMP subtests shows a large, significant effect of the symbol-retrieval mechanism (χ2 (6) = 134.29, p < 0.001). Contrasts against the learned-relative-symbol baseline (Table 12) show that the two relationalsymbolic conditions are significantly worse on BLiMP than the learned-relative baseline’s 69.34%: relsymbolic_n4, which combines the larger size4 relational-symbol neighborhood with a learned symbol library, at 66.32% (p < 0.001), and relsymbolic at 68.33% (p = 0.023). The remaining five conditions (learned-relative, RoPErelative, learned-positional, sinusoidal-positional,

and symbolic-attention) are not significantly different from one another (all pairwise p > 0.5, Tukeyadjusted); beyond the relational-symbolic disadvantage, the comparison does not support a finer ranking among symbol sources. That indifference is itself useful: RoPE-based relative symbols match the learned library they replace (69.43% vs. 69.34%) while eliminating its parameter cost—a learned relative-symbol library holds (2∆+1) d parameters, 1.05M for the submitted 1,024-dimensional models with ∆=512—and its maximum-offset table. 4.5

Which Relational Attention Type?

The controlled RA-type comparison varies only attention type, objective, and LM-head tying across 25 runs, holding architecture, symbols, context, and recipe fixed (Appendix E). On BLiMP, the three types are indistinguishable: the internal-DAT GLMM finds no significant effect of attention type (χ2 (2) = 3.50, p = 0.174), and the largest between-type spread in any condition is 0.76 points against a pooled within-cell seed standard deviation of 0.54. RCA is numerically highest in all three conditions. This converges with the DAT paper’s relational-games ablation, where RA and RCA perform similarly (Altabaa and Lafferty, 2025), and contradicts its language-modeling ablation, in which RCA-head DAT performs “no better than a standard Transformer with a matching total number of heads”: ours beats exactly that baseline by 2.5 BLiMP points (Section 4.2). At 10M words the simpler RCA, which routes only symbols and needs no relation-projection parameters, is the equal of full RA. At 100M words the ordering changes. In a matched 18-layer pair differing only in attention type, RA beats RCA on six of seven benchmarks (BLiMP 79.30 vs. 78.44, EWoK +1.7, entity tracking +3.7). RA is also the least stable: in a controlled 12-layer, 1,024-dimensional pair its training loss diverged and never recovered below its epoch0 value, while RCA trained monotonically to 79.72 BLiMP and DisRCA to 79.38 (Appendix E). The extra capacity RA carries—separate relation projections and four relation channels—appears to need both data and a stable recipe. 4.6

SA/RA Head-Ratio Sweep

We sweep the sensory-to-relational head ratio at 512-token context, with 3 seeds per ratio (Table 18, Appendix E). A mixed-effects model with

seed as a random intercept shows that higher relational-attention (RA) fraction significantly decreases BLiMP accuracy (slope = −1.827, t = −5.887, p < 0.001) and EWoK accuracy (slope = −1.039, p = 0.021). With 12 heads in total the two fractions are complements: a larger relational stream is also a smaller sensory one. The RA fraction has no significant effect on COMPS (p = 0.201) or reading measures (p = 0.488), and COMPS and supplement scores stay flat across ratios, so the head-ratio allocation acts mainly on syntactic structures. Mean BLiMP peaks at the balanced 6SA/6RA split (70.33%), but 9SA/3RA — the allocation used in all submitted models — trades 0.6 BLiMP points for a 1.4-point supplement gain and the best COMPS score.

5

Related Work

BabyLM and data-efficient language models. The BabyLM challenge (Warstadt et al., 2023; Choshen et al., 2026) provides a shared evaluation for sample-efficient pretraining on developmentally plausible corpora. Successful entries have mostly intervened on the data or the training objective rather than on the attention mechanism: curricula and corpus filtering, distillation, and hybrid objectives such as GPT-BERT (Charpentier and Samuel, 2024), which merges causal and masked prediction in one Transformer stack. Architecture-level interventions like our submission evaluated under the full suite are comparatively rare. Relational bottleneck and dot-product relations. The closest point of comparison to our work is the Abstractor architecture and relational crossattention (RCA) (Altabaa et al., 2024) and DAT (Altabaa and Lafferty, 2025). While DAT greatly improves data efficiency on purely relational tasks, the evidence for its usefulness in language modeling is a modest perplexity gain on web text (16.94 to 16.09 at the 350M scale, and less at larger scales), with no evaluation of what the relational stream does to linguistic generalization. Our work extends the original DAT design with new symbol-retrieval mechanisms and studies them in a developmentally plausible training regime with controlled comparisons under the BabyLM evaluation suite. The motivation for RCA and DAT connects to older work on compositionality and variable binding (Smolensky, 1990; Holyoak and Hummel, 2000). The ESBN (Webb et al., 2021) shows that a recurrent network augmented with an external

memory can implement variable binding and indirection, allowing symbol-like representations to emerge and enabling near-perfect generalization of abstract rules to novel entities; CoRelNet (Kerg et al., 2022) simplifies this mechanism, showing that a matrix of dot-product similarities can itself serve as a relational representation that supports out-of-distribution generalization. A broader cognitive framework for all of these works is provided in the relational bottleneck proposal (Webb et al., 2024), which argues that architectures can benefit from restricting part of the computation to relations rather than object-specific sensory content. Latent-state objectives and cognitive/psycholinguistic evaluation. Next-Latent Prediction (Teoh et al., 2026) was introduced and evaluated as a world-modeling objective, on world-model benchmarks and on the latent rollouts that support speculative decoding. Whether the compression pressure it imposes also makes a model more human-like is a separate question, and the one our reading regression asks (Section 4.3). It is the same pressure the Now-or-Never bottleneck attributes to human comprehension (Christiansen and Chater, 2016), and the link between a model’s word-by-word predictions and human reading effort is the standard currency for testing it (Hale, 2001; Lewis and Vasishth, 2005).

6

Conclusion

We set out to ask whether the data efficiency shown by relational attention on purely relational tasks transfers to language modeling. Our answer is a qualified yes. Under the strict-small regime, architecture is the dominant factor for structural linguistic generalization, worth roughly 2.5 BLiMP points between the best configurations of each at matched depth, width, and head count; the gain comes from adding a relational stream rather than from enlarging it, and at this scale the three relational mechanisms are interchangeable. The NextLat objective is secondary but significant, and interacts with architecture: NextLat models track human reading measures more closely and improve 5 of 7 finetuned (Super)GLUE accuracies. Among symbol sources only the relational-symbolic variants underperform, and our novel parameter-free RoPE-based relative symbols match the learned library they replace. At 100M words full relational attention pulls ahead, and our two strongest models rank among the top strict-track entries at the time of writing.

Limitations Statistical caveats. The binomial GLMMs model aggregated per-item counts without observation-level random effects, so overdispersion is not addressed and exact p-values may be anti-conservative. Incomplete replication. Most experiment groups contain only 3 seeds; only the RCA/SwiGLU 264-token symbol-retrieval comparison has the full 5-seed design. NextLat effect size. While the training objective is statistically significant for cognitive alignment (t = 5.22, p < 0.001 in the reading regression), the effect is secondary to the architecture effect on structural generalization. The architecture×objective interaction (χ2 (1) = 10.81, p = 0.001) indicates that the NextLat benefit is not uniform across architectures. The evidence for the NextLat inductive bias driving structural generalization, while statistically significant, does not have the same weight as architecture in our experiments. Scope. All replicated results use the strict-small (10M-word) track. Our submitted strict-track entries are single-seed; the recipe-comparison models have two (NextLat) and three (NTP) seeds. At strict scale the base and NextLat runs also differ in LMhead tying and parameter count; the strict-small matrix separates the two, since its untied/NTP and untied/NextLat cells are matched on both. Seed spread is non-trivial: the Muon/LambW NTP trio spans 79.81/79.20/78.78 BLiMP (range 1.03) and 56.66/55.23/55.35 EWoK (range 1.43); the NextLat pair spans 79.75/79.03 BLiMP (range 0.72). The seed ranges exceed the NTP-vs-NextLat gaps the recipe table reports (0.06 BLiMP, 0.51 EWoK on seed 0). Entity tracking has one extreme outlier: a third NTP seed scores 33.63 against 20.84 and 20.80 for the other two. The comparison with the official GPT-2 baseline is likewise not parameteror recipe-matched. We do not evaluate on the strict (100M-word) track with full replication, leaving open whether the DAT architecture and NextLat objective effects scale with more data. The standard Transformer baseline uses 12 self-attention heads; we do not report a parameter-matched baseline that scales head count independently of the sensory/relational split. The architecture comparison is also not matched on every other axis: the DAT and standard configurations differ in feed-forward

activation, initialization, and symbol mechanism as well as in attention. Dataset and task coverage. We evaluate on the BabyLM strict-small zero-shot suite (BLiMP, BLiMP supplement, COMPS, Entity Tracking, EWoK, and the human-likeness measures) and the strict-track (Super)GLUE fine-tuning suite. We do not report generation quality. The human-likeness measures derive from a single resource of 205 sentences (de Varda et al., 2023) covering eye-tracking, self-paced reading, and EEG responses. Implementation faithfulness. The DAT implementation in our experiments differs from the original DAT formulation in several respects: the use of SwiGLU vs. GELU feed-forward activations, tied vs. untied LM heads, the symbol-retrieval mechanism (RoPE-based relative symbols rather than a learned symbol library), and the weight initialization scheme. The relation activation used in RCA (identity or sigmoid rather than softmax) is a hyperparameter setting the original design provides for (in principle), but has remained previously underexplored. Optimizer-recipe comparison. The Muon/LambW recipe improves all five accuracybased zero-shot measures (BLiMP +1.73pp, BLiMP supplement +1.15pp, EWoK +1.77pp, entity tracking +1.99pp, COMPS +1.02pp; Table 7), while reading-time measures decrease slightly. This comparison jointly changes optimizer family, learning rate, and weight decay, so the improvement is a property of the recipe and (potentially) not of specific components in isolation.

Acknowledgments AB gratefully acknowledges the compute support from the National Research Platform (NRP) at the University of California, San Diego. NRP has been developed, and is supported in part, by funding from National Science Foundation, from awards 1730158, 1540112, 1541349, 1826967, 2112167, 2100237, and 2120019, as well as additional funding from community partners (Weitzel et al., 2025). ET and JD were supported by the European Research Council (ERC), grant 101088098 - MEMLANG. Views and opinions expressed are however those of the authors only and do not necessarily reflect those of any funding agency. Part of this work also used resources available through the Dutch

national e-infrastructure with the support of the SURF Cooperative using grant no. EINF-18356. LLM assistants were used in circumscribed, closely supervised ways during code development and the editing process of the manuscript. On the spectrum between local, narrow autocomplete vs. global, large-scale LLM-based generation of code or text, this work stays decidedly close to the local / narrow end of the spectrum. We retain full responsibility for all content, and all errors are our own.

References Awni Altabaa and John Lafferty. 2025. Disentangling and integrating relational and sensory information in transformer architectures. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 1271–1297. PMLR. Awni Altabaa, Taylor Whittington Webb, Jonathan D. Cohen, and John Lafferty. 2024. Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in transformers. In The Twelfth International Conference on Learning Representations. Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah, Abdelrahman Eldesokey, Abeer Kashar, Abolade Daud, Abosede Grace Olanihun, Adamu Labaran Mohammed, Adeyemi Praise, Adhikarimayum Meerajita Sharma, and 1 others. 2025. Global PIQA: Evaluating commonsense reasoning across 100+ languages and cultures. Preprint, arXiv:2510.24081. Tyler A. Chang and Benjamin K. Bergen. 2022. Word acquisition in neural language models. Transactions of the Association for Computational Linguistics, 10:1–16. Lucas Georges Gabriel Charpentier and David Samuel. 2024. GPT or BERT: why not both? In The 2nd BabyLM Challenge at the 28th Conference on Computational Natural Language Learning, pages 262– 283, Miami, FL, USA. Association for Computational Linguistics. Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Jaap Jumelet, Tal Linzen, Aaron Mueller, Suchir Salhan, Raj Sanjay Shah, Alex Warstadt, and Ethan Gotlieb Wilcox. 2026. BabyLM turns 4 and goes multilingual: Call for papers for the 2026 BabyLM workshop. Preprint, arXiv:2602.20092. Morten H. Christiansen and Nick Chater. 2016. The now-or-never bottleneck: A fundamental constraint on language. Behavioral and Brain Sciences, 39:e62. Andrea Gregor de Varda, Marco Marelli, and Simona Amenta. 2023. Cloze probability, predictability ratings, and computational estimates for 205 English

sentences, aligned with existing EEG and reading time data. Behavior Research Methods, 56(5):5190– 5213. David R. Dowty, Robert E. Wall, and Stanley Peters. 1981. Introduction to Montague Semantics. Studies in Linguistics and Philosophy. D. Reidel Publishing Company, Dordrecht. Jeroen Groenendijk and Martin Stokhof. 1991. Dynamic predicate logic. Linguistics and Philosophy, 14(1):39–100. John Hale. 2001. A probabilistic Earley parser as a psycholinguistic model. In Second Meeting of the North American Chapter of the Association for Computational Linguistics. John T. Hale. 2011. What a rational parser would do. Cognitive Science, 35(3):399–443. Irene Heim. 1982. The Semantics of Definite and Indefinite Noun Phrases. Ph.D. thesis, University of Massachusetts Amherst, Amherst, MA. Keith J. Holyoak and John E. Hummel. 2000. The proper treatment of symbols in a connectionist architecture. In Eric Dietrich and Arthur B. Markman, editors, Cognitive Dynamics: Conceptual and Representational Change in Humans and Machines, pages 229–263. Lawrence Erlbaum Associates, Mahwah, NJ. Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas H. Clark, Carina Kauf, Jennifer Hu, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan G. Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Joshua Tenenbaum, and Jacob Andreas. 2025. Elements of world knowledge (EWoK): A cognition-inspired framework for evaluating basic world knowledge in language models. Transactions of the Association for Computational Linguistics, 13:1245–1270. Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. 2024. Muon: An optimizer for hidden layers in neural networks. Hans Kamp. 1981. A theory of truth and semantic representation. In Jeroen Groenendijk, Theo Janssen, and Martin Stokhof, editors, Formal Methods in the Study of Language, pages 277–322. Mathematical Centre Tracts, Amsterdam. Giancarlo Kerg, Sarthak Mittal, David Rolnick, Yoshua Bengio, Blake Richards, and Guillaume Lajoie. 2022. On neural architecture inductive biases for relational tasks. Preprint, arXiv:2206.05056. Najoung Kim and Sebastian Schuster. 2023. Entity tracking in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3835–3855, Toronto, Canada. Association for Computational Linguistics.

Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations. Richard L. Lewis and Shravan Vasishth. 2005. An activation-based model of sentence processing as skilled memory retrieval. Cognitive Science, 29(3):375–419. Jingyuan Liu, Jianlin Su, Xingcheng Yao, Zhejun Jiang, Guokun Lai, Yulun Du, Yidao Qin, Weixin Xu, Enzhe Lu, Junjie Yan, Yanru Chen, Huabin Zheng, Yibo Liu, Shaowei Liu, Bohong Yin, Weiran He, Han Zhu, Yuzhi Wang, Jianzhou Wang, and 9 others. 2025. Muon is scalable for LLM training. Preprint, arXiv:2502.16982. Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations. Gary F. Marcus, Sujith Vijayan, Shoba Bandi Rao, and Peter M. Vishton. 1999. Rule learning by sevenmonth-old infants. Science, 283(5398):77–80.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008. Curran Associates, Inc. Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjape, Adina Williams, Tal Linzen, and Ryan Cotterell. 2023. Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora. In Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, pages 1–34, Singapore. Association for Computational Linguistics. Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. BLiMP: The benchmark of linguistic minimal pairs for English. Transactions of the Association for Computational Linguistics, 8:377– 392.

Kanishka Misra, Julia Rayz, and Allyson Ettinger. 2023. COMPS: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2928– 2949, Dubrovnik, Croatia. Association for Computational Linguistics.

Taylor W. Webb, Steven M. Frankland, Awni Altabaa, Simon Segert, Kamesh Krishnamurthy, Declan Campbell, Jacob Russin, Tyler Giallanza, Randall O’Reilly, John Lafferty, and Jonathan D. Cohen. 2024. The relational bottleneck as an inductive bias for efficient abstraction. Trends in Cognitive Sciences, 28(9):829–843.

Richard Montague. 1973. The proper treatment of quantification in ordinary english. In Approaches to Natural Language, pages 221–242. D. Reidel Publishing Company, Dordrecht.

Taylor Whittington Webb, Ishan Sinha, and Jonathan D. Cohen. 2021. Emergent symbols through binding in external memory. In The Ninth International Conference on Learning Representations.

Barbara H. Partee. 1995. Lexical semantics and compositionality. In Lila R. Gleitman and Mark Liberman, editors, An Invitation to Cognitive Science, Volume 1: Language, 2 edition, pages 311–360. MIT Press, Cambridge, MA.

Derek Weitzel, Ashton Graves, Sam Albin, Huijun Zhu, Frank Wuerthwein, Mahidhar Tatineni, Dmitry Mishin, Elham Khoda, Mohammad Sada, Larry Smarr, Thomas DeFanti, and John Graham. 2025. The national research platform: Stretched, multitenant, scientific Kubernetes cluster. In Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration, PEARC ’25, New York, NY, USA. Association for Computing Machinery.

Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. Technical report, OpenAI. Paul Smolensky. 1990. Tensor product variable binding and the representation of symbolic structures in connectionist systems. Artificial Intelligence, 46(1– 2):159–216. Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. Roformer: Enhanced transformer with rotary position embedding. Preprint, arXiv:2104.09864. Originally published in 2021. Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Tim Pearce, Pratyusha Sharma, Akshay Krishnamurthy, Riashat Islam, Alex Lamb, and John Langford. 2026. Next-latent prediction transformers learn compact world models. Preprint, arXiv:2511.05963. Version 3, May 2026; v1 posted November 2025.

Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020. Large batch optimization for deep learning: Training BERT in 76 minutes. In International Conference on Learning Representations.

A

Further Details

Our code—model implementation, training and evaluation scripts, configurations, and analysis code—is available at https://github.com/ abrsvn/babylm_dat_2026. Our submitted checkpoints are also publicly available on the HuggingFace Hub and are linked from the official leaderboard entries.

Architecture

Symbols Head ratio

Context 512 264 512 Batch size 12 or 16 16 12 LR (AdamW) 5 × 10−5 10−3 5 × 10−5 SA/RA split 6/6, 9/3 9/3 1/11 to 11/1 Attn type RA, RCA, DisRCA RCA RCA Rel. activation identity sigmoid identity Activation GELU/SwiGLU SwiGLU SwiGLU LM head tied/untied tied tied Seeds 3 5 3

Table 3: Training configurations of the three strictsmall analysis families (Section 3.4). SA/RA = sensory/relational head split. All runs use AdamW, 10 epochs, gradient clipping at 1.0, and 12-layer, 768dimensional backbones. The architecture comparison also spans training objective (NTP vs. NextLat) and includes single auxiliary runs at 264- and 1,280-token context (the context-length factor in the internal-DAT analysis of Appendix E); LM-head tying varies independently of the objective: NextLat runs are untied, NTP runs appear both ways.

B

Strict-Small Training Configurations

Table 3 summarizes the training configurations of the three strict-small analysis families (Section 3.4), referenced from Section 3.2.

C

Top Strict-Track Model Specifications

Table 4 gives the full model and training specifications of our two strongest strict-track entries (Section 4.1). Parameter counts are computed directly from the published checkpoints; the remaining settings are taken from the training configurations. Both models use the RoPE-based relative symbols introduced in Section 2.2, which add no learned parameters, and both submitted checkpoints are exponential moving averages of the training weights. Table 5 reports the official per-task fine-tuned (Super)GLUE scores for the two top entries and the GPT-2 baseline.

D

Benchmark Definitions

BLiMP. The Benchmark of Linguistic Minimal Pairs (Warstadt et al., 2020) consists of 67 minimalpair paradigms testing grammatical phenomena such as subject–verb agreement, island effects, and filler–gap dependencies. We use aggregate BLiMP accuracy and a decomposition by linguistic category. BLiMP supplement. The BLiMP supplement contains additional minimal-pair paradigms be-

yond the core 67, probing further syntactic and semantic phenomena. COMPS. The Conceptual Minimal Pair Sentences benchmark (Misra et al., 2023) tests property attribution and property inheritance. Each pair attributes a property to two concepts, one that has it and one that does not (A robin/#penguin can fly); the harder subsets restate the property of a novel concept introduced as a subordinate of the positive or negative concept (A wug is a robin/penguin. Therefore, a wug can fly), which controls for memorization of the literal test phrase, and then add the negative concept back as a distractor. Entity Tracking. This benchmark (Kim and Schuster, 2023) evaluates whether a model can maintain and update the changing states and locations of entities across a discourse, testing a core cognitive prerequisite for narrative comprehension. EWoK. The Elements of World Knowledge benchmark (Ivanova et al., 2025) evaluates basic conceptual world knowledge across domains such as spatial relations, physical states, and social interactions. Each example consists of minimal-pair sentences testing whether a model understands concepts that help model the world. Reading and eye-tracking. The BabyLM human-likeness benchmark regresses a model’s per-word surprisal against eleven psycholinguistic dependent variables drawn from de Varda et al. (2023): four eye-tracking measures (first-fixation, first-pass, go-past, and right-bounded reading time), self-paced reading time, and six eventrelated potential components (ELAN, LAN, N400, P600, EPNP, PNP). The incremental R2 (∆R2 ) reported in Table 9 is the additional variance in these measures explained by adding model surprisal to baseline predictors, averaged over the eleven variables. GlobalPIQA. Global PIQA (Chang et al., 2025) is a physical commonsense-reasoning benchmark spanning over 100 language varieties, with a culture-specific non-parallel split and a shared, culture-agnostic parallel split. The Strict and Strict-Small tracks use the English subset, evaluated cloze-style: completions are chosen by length-normalized conditional log-probabilities; the leaderboard score is the mean of the parallel and non-parallel splits (Choshen et al., 2026).

Setting

18L curriculum model

Leaderboard name Parameters Layers / hidden dim. Attention heads (SA/RA) Attention type Symbol mechanism

DAT Strict Curriculum DAT Strict NextLat Final 191.2M 304.3M 18 / 768 16 / 1,024 9/3 12 / 4 RCA RA (4 relation channels, identity activation) RoPE-based relative symbols, shared across layers; no learned symbol library (parameterfree) SwiGLU, hidden dim. 3,072 SwiGLU, hidden dim. 4,096 pre-LayerNorm RoPE (θ=10,000); max. relative position 512 untied 16,384 / 512 packed tokens (514 model positions) GPT-BERT 100M (2024 release) GPT-BERT 100M (2026 release) 48 sequences of 512 tokens 36 sequences of 512 tokens 1 1 Phase A: 6 epochs line-EOS; Phase B: 4 Phase A: 2 epochs line-EOS; Phase B: 8 epochs document-boundary epochs document-boundary 0.02 / 0.008 0.02 / 0.005 10−3 / 4 × 10−4 10−3 / 2.5 × 10−4 0.01 0.05 1% warmup, cosine decay to 10% of the initial rate horizon 1; MSE weight 1.0, KL weight 0.5, CE weight 0.0 decay 0.999; submitted weights are the EMA checkpoints scaled-normal (0.02) projections / 0.1

Feed-forward Normalization Positional encoding LM head Vocabulary / sequence length Tokenizer Batch Random seed Curriculum Muon lr (Phase A / Phase B) LambW lr (Phase A / Phase B) Weight decay LR schedule NextLat EMA Initialization / dropout

16L wide model

Table 4: Full specifications of our two strongest strict-track models, verified against the published checkpoints and training configurations. Both are decoder-only DAT models trained from scratch on the BabyLM strict (100M-word) corpus for 10 total epochs: Muon (momentum 0.95, 5 Newton–Schulz steps) optimizes hidden weight matrices and LambW (β1 =0.9, β2 =0.999, ϵ=10−8 ) optimizes embeddings, heads, symbol tables, norms, and biases. Phase B is warm-started from the Phase-A EMA weights. The two models use different releases of the GPT-BERT 100M-word tokenizer, both with a vocabulary of 16,384: the 2024 release is byte-identical to the tokenizer distributed with the official GPT-BERT baseline, and the 2026 release is a retraining of the same tokenizer recipe on the 2026 strict corpus (Section 3.1). Training sequences are packed to 512 tokens as in Section 3.1; the models allocate 514 positions: the 512 packed tokens, a leading BOS slot, and one appended final-target token that the one-step NextLat rollout conditions on (nextlat-free models allocate 513). Parameter counts are checkpoint (backbone) counts; the training-time NextLat dynamics module adds a further 10.5M parameters and is discarded after training. No teacher model, synthetic data, or multimodal input is used.

Age of acquisition (AoA). Following Chang and Bergen (2022), the AoA evaluation tracks word surprisal across training checkpoints, derives a model age of acquisition for each word, and compares it with children’s acquisition ages from CDI norms. The reported score is the Pearson correlation between model and child acquisition ages (scaled by 100); correlations that are not significant at p ≤ 0.1 are scored as zero.

E

Further Results

Strict-small leaderboard results. Table 6 reports the official strict-small (10M-word) leaderboard results for our two submissions and the GPT2 baseline, complementing the strict-track results of Table 1 and discussed in Section 4.1. Optimizer-recipe comparison. Table 7 reports the strict-track optimizer-recipe comparison between AdamW and Muon/LambW at 12 layers,

referenced in Sections 3.3 and 4.1. Strict-track fine-tuned task performance. Table 8 reports the fine-tuned (Super)GLUE comparison between the same two 12-layer Muon/LambW models as Table 7, discussed in Section 4.3. Reading regression. Table 9 gives the fixedeffect coefficients of the cross-experiment reading regression of Sections 3.5 and 4.3. Architecture-comparison group means. Table 10 gives the per-configuration means behind Section 4.2. Symbol ablation at submission scale (not controlled). A strict-small run with a learned relative-symbol library in place of RoPE-based symbols scores 71.42 on BLiMP against 72.33 for the submitted model, but is 14 layers deep rather than 12 and was trained with an exponential mov-

Task

GPT-2 base

18L curric.

16L wide

BoolQ MNLI MRPC MultiRC QQP RTE WSC

69.66 60.76 85.34 65.92 71.56 57.55 63.46

69.48 63.00 88.00 65.68 73.14 66.91 63.46

70.52 66.26 88.74 66.30 74.26 70.50 63.46

(Super)GLUE avg.

67.75

69.95

71.43

Table 5: Official fine-tuned (Super)GLUE task scores for the GPT-2 baseline and our two top strict-track entries (Table 1), from the official leaderboard evaluation. Metrics follow the leaderboard’s definitions: F1 for MRPC and QQP, accuracy for the rest—including MultiRC, which the BabyLM pipeline scores by accuracy rather than by SuperGLUE’s F1a and exact match. Bold marks the best value per row.

Figure 2: Mean BLiMP accuracy across model configurations. With few exceptions, DAT architectures outperform the standard Transformer baseline (top). Error bars are ±1 standard error across seeds.

ing average, so the difference is not attributable to the symbol mechanism. A depth-matched 12-layer run completed training but was not evaluated. Attention type at 100M words (18 layers). Two 18-layer curriculum runs differ only in attention type and in batch size (48 for RCA, 44 for RA; the RA model’s relation projections raise its memory cost): same 2026 tokenizer, same phase-A switch point, same learning rates, weight decay, and epoch counts, each warm-started from its own phase A. RA scores higher on six of seven benchmarks — BLiMP 79.30 vs. 78.44, EWoK 57.93 vs. 56.20, entity tracking 22.77 vs. 19.07, COMPS 58.62 vs. 57.23, and both GlobalPIQA splits — and lower on the BLiMP supplement (68.56 vs. 69.00). Each arm is a single seed, and RA is sensitive to the surrounding hyperparameters: two further 18-layer RA runs under different learning-rate, switch-point, and epoch settings score 78.18 and 76.49, a range

that spans the RCA result. All figures in this paragraph are raw-checkpoint scores, not exponential moving averages. Wide-12L training stability (controlled for RA vs. RCA). At 1,024 dimensions and 12 layers, a pair differing only in attention type shows RA diverging during training (loss peaking at 5.25 in epoch 3, never recovering below its epoch-0 value; BLiMP 58.21, EWoK 50.15, COMPS 49.65, all at or near chance) while RCA trains monotonically (BLiMP 79.72). A DisRCA run also trains cleanly (BLiMP 79.38) but additionally differs in weight decay (0.01 vs. 0.05). The submitted 16-layer wide RA model trained normally under a two-phase curriculum with a reduced Phase-B learning rate. Perepoch training loss: RA 3.44, 3.69, 4.55, 5.25, 5.14, 4.97, 4.49, 3.98, 3.75, 3.61; RCA 3.34, 2.81, 2.67, 2.40, 2.36, 2.33, declining monotonically; DisRCA 3.36, 2.95, 2.90, 2.87, 2.84, 2.82, 2.78, 2.74, 2.70, 2.68. Tokenizer and curriculum observations (not controlled). Two strict-track observations from the 18-layer curriculum experiments are worth recording, though neither is a controlled comparison. First, the 2024-tokenizer run scores slightly higher than the 2026 retraining on the same 18-layer architecture (BLiMP EMA 79.17 vs. 78.44; supplement 70.45 vs. 69.00), but the runs also differ in weight decay (0.01 vs. 0.05), phase-A switch point, and Phase-B epoch count, so the tokenizer effect cannot be isolated. Second, two 2026-tokenizer RA curriculum runs with different hyperparameter settings (learning rate 4e-4 vs. 1e-3, Muon lr 0.008 vs. 0.02, phase-A switch at checkpoint 3 vs. 1, PhaseB epochs 6 vs. 8) score BLiMP 79.30 and 76.49; the lower-learning-rate, later-switch setting wins. A curriculum-versus-document-boundary pair is confounded by warm-start, learning rate, Muon lr, and epoch count (weight decay is identical). Baseline-vs-DAT omnibus tests. Table 11 reports the Type III Wald χ2 tests for the baseline-vsDAT GLMM of Section 4.2, and Table 12 the full pairwise contrasts for the symbol-retrieval comparison of Section 4.4. Internal-DAT analysis. The internal-DAT ANOVA (Table 13) decomposes BLiMP performance across the DAT design factors. The activation/symbol-position term (χ2 (2) = 26.41, p < 0.001) covers the early GELU/symbolic runs that preceded the fixed-SwiGLU design

BLiMP Supp. EWoK Ent. trk. COMPS PIQA GLUE∗ Reading

Model

AoA Overall

GPT-2 baseline (12L)

65.23 57.25

50.63

19.10

51.81 35.09

63.80

5.63 −12.15

37.38

DAT 12L (NTP) DAT 12L (NextLat)

70.40 57.94 72.34 56.40

50.25 52.93

18.62 18.45

52.79 35.64 52.69 33.65

62.96 63.97

6.56 7.12

−14.00 −14.15

37.91 38.16

Table 6: Strict-small (10M-word) results from the official BabyLM leaderboard at the time of writing (127 entries). Both DAT models use 9SA/3RA RCA, SwiGLU activations, RoPE-based relative symbols, and the Muon/LambW recipe (Muon lr 0.02, effective batch of 16 sequences of 512 tokens, seed 1, 10 epochs; the NextLat run reaches that batch as 8 sequences with 2 gradient-accumulation steps); the NTP model has a tied LM head (123M parameters) and the NextLat model an untied head (136M parameters). Column definitions as in Table 1. Bold marks the best value per column. The comparison with the GPT-2 baseline (98M parameters, AdamW) is not parameter- or recipe-matched. Optimizer

Objective

BLiMP

Supp.

EWoK

Ent. trk.

COMPS

AdamW AdamW Muon/LambW Muon/LambW

NTP NextLat NTP NextLat

78.08 77.95 79.81 79.75

58.88 59.92 60.03 61.01

54.89 55.07 56.66 57.17

18.85 19.63 20.84 20.27

56.82 57.69 57.84 57.90

Table 7: Optimizer-recipe comparison on the strict track: 12-layer, 768-dimensional DAT models (9SA/3RA, SwiGLU, relative symbols), batch of 64 sequences. AdamW: lr 5 × 10−4 , no weight decay. Muon/LambW: Muon lr 0.02 (momentum 0.95, 5 Newton–Schulz steps) for hidden matrix parameters; LambW lr 10−3 (weight decay 0.01) for embeddings, heads, norms, and biases. All rows are seed 0. NTP models use tied LM heads; NextLat models use untied LM heads. The Muon/LambW recipe improves all five measures, but the comparison jointly changes optimizer, learning rate, and weight decay.

(Section 3.4); LM-head tying (p = 0.006) and context length (p = 0.001) are also significant. Relational attention type (p = 0.174) and layer allocation (p = 0.437) are not (Section 4.5). The training objective is also not significant within the DAT family on BLiMP (p = 0.216), consistent with architecture being the primary lever for structural generalization while NextLat provides a complementary cognitive-alignment benefit. Objective at matched LM-head tying. The architecture comparison observes three of the four head-tying × objective combinations: tied/NTP, untied/NTP, and untied/NextLat. The last two are matched on head tying and on parameter count, since the auxiliary dynamics model is discarded after training, so their difference isolates the objective. Within the controlled cell of Table 14, NextLat gains 0.83 BLiMP points over NTP (RCA +1.17, DisRCA +0.67, RA +0.58); the internal-DAT model does not find the objective significant on BLiMP (p = 0.216), while it is highly significant in the reading regression (t = 5.22, p < 0.001). The first two cells isolate LM-head tying at NTP: 70.18 tied against 69.42 untied. RA-type per-cell means. Table 14 reports BLiMP means for the controlled RA-type com-

parison of Section 4.5: 25 runs holding SwiGLU, learned relative symbols, 512-token context, 9SA/3RA, identity relation activation, 12 layers, 768 dimensions, and all training hyperparameters fixed. Three of the seven RA runs used a batch of 16; the other 22 used 12. Pooled within-cell seed SD is 0.54 (df=16). Further symbol-retrieval results. Table 15 gives the per-condition means behind Section 4.4, and Table 16 the same comparison on composite reading scores. SA/RA head-ratio sweep. Figure 3 shows mean BLiMP accuracy across SA/RA head-ratio splits, and Table 19 reports the mixed-effects model estimates for the SA/RA head-ratio sweep at 512token context (3 seeds per ratio, seed as random intercept). The RA head fraction has a significant negative effect on BLiMP and EWoK, but no significant effect on COMPS or reading measures.

Task BoolQ MNLI MRPC MultiRC QQP RTE WSC

Muon base

Muon NL

∆

0.684 0.596 0.725 0.598 0.761 0.547 0.654

0.684 0.594 0.740 0.655 0.772 0.576 0.692

0.000 −0.002 +0.015 +0.057 +0.011 +0.029 +0.038

Table 8: Fine-tuned (Super)GLUE task accuracies: Muon/LambW base (no NextLat, tied head) vs. Muon/LambW NextLat (untied head). All seven tasks are scored by accuracy (MultiRC by per-answer accuracy, rather than SuperGLUE’s F1a and exact match); the leaderboard itself scores MRPC and QQP by F1, so these values differ from the corresponding leaderboard entries for the same checkpoints. NextLat improves 5 of 7 task accuracies; a two-sided paired t-test across the seven task accuracies gives t(6) = 2.61, p = 0.040. These two strict-track models differ in LM-head tying and parameter count (≈ 12.6M); the strict-small matrix separates the objective from tying, since its untied/NTP and untied/NextLat cells are matched on both.

Term

Estimate

SE

t

p

(Intercept) model_type objective tie_lm_head datapoint_length1 datapoint_length2 model_variant

0.048 −0.001 0.002 0.000 −0.005 −0.004 −0.003

0.011 0.000 0.000 0.000 0.001 0.001 0.000

4.56 −2.61 5.22 0.35 −7.08 −4.85 −28.48

0.001 0.010 < 0.001 0.728 < 0.001 < 0.001 < 0.001

Table 9: Cross-experiment reading regression: coefficients for the incremental R2 (∆R2 ) gained by adding model surprisal to baseline reading-time predictors. The response is the per-experiment, per-dependent-variable ∆R2 . The training objective is significant (t = 5.22, p < 0.001): Next-Latent Prediction models explain more variance in human reading measures than NTP models. Model type (DAT vs. standard) is also significant (p = 0.010). The model-variant effect (p < 0.001) reflects the large difference between base and spillover reading conditions.

Figure 3: Mean BLiMP accuracy across SA/RA headratio splits. Error bars are ±1 standard error across 3 seeds.

Architecture

Objective

LM head

N

BLiMP

Reading

Entity

COMPS

Standard Transformer Standard Transformer Standard Transformer

NTP NTP NextLat

tied untied untied

3 3 3

67.76 68.16 67.68

7.06 7.16 7.36

14.27 19.14 16.64

52.10 51.86 52.03

DAT (RCA, SwiGLU) DAT (RCA, SwiGLU) DAT (RA, GELU) DAT (RA, GELU)

NTP NextLat NTP NextLat

tied untied tied untied

3 3 3 3

70.66 70.62 70.04 69.39

6.66 7.27 6.89 7.37

19.71 21.92 22.23 20.46

52.42 52.36 52.51 52.28

Table 10: Main results: BLiMP accuracy and downstream benchmarks by architecture and training objective. Standard Transformer baselines and DAT models share the same 12-layer, 768-dimensional, 12-head configuration, but also differ in feedforward activation (GELU vs. SwiGLU), initialization, and symbol mechanism. N is the number of random seeds. Reading is the composite reading score (mean of eye-tracking and self-paced reading).

Effect

Effect (Intercept) base_arch is_tied objective base_arch:is_tied base_arch:objective

χ2

df

p

39.76 139.43 2.64 2.26 10.62 10.81

1 1 1 1 1 1

< 0.001 < 0.001 0.105 0.132 0.001 0.001

(Intercept) attn_type layers activation_position is_tied objective context_length

Table 11: Type III ANOVA (Wald χ2 tests) for the baseline-vs-DAT BLiMP model (binomial GLMM). The model includes fixed effects for base architecture (standard Transformer vs. DAT), LM-head tying, and training objective, with crossed random intercepts for experiment, BLiMP subtest, and pseudo-item. The basearchitecture effect is large and significant; the architecture also interacts significantly with both head tying and objective.

χ2

df

p

56.999 3.496 0.604 26.406 7.564 1.531 10.187

1 2 1 2 1 1 1

< 0.001 0.174 0.437 < 0.001 0.006 0.216 0.001

Table 13: Type III ANOVA (Wald χ2 tests) for internal DAT mechanics on BLiMP (binomial GLMM). Fixed effects: relational attention type (RA/RCA/DisRCA), layer allocation, activation/position setting, LM-head tying, training objective, and context length. Activation/position, head tying, and context length are significant; relational attention type and layer allocation are not. Condition

RA

RCA DisRCA spread

tied / NTP 69.99 (3) 70.66 (3) 69.90 (3) untied / NTP 69.36 (2) 69.46 (3) 69.41 (3) untied / NextLat 69.94 (2) 70.62 (3) 70.08 (3)

0.76 0.10 0.68

marginal

0.45

69.79 (7) 70.25 (9) 69.80 (9)

Table 14: BLiMP cell means for the controlled RA-type comparison. RCA is numerically highest in all three conditions; the omnibus GLMM finds no significant effect of attention type (χ2 (2) = 3.50, p = 0.174). Contrast

Estimate

SE

z

p

relative − relative_rope relative − positional relative − positional_sinusoidal relative − symbolic relative − relsymbolic relative − relsymbolic_n4

−0.006 0.012 0.030 0.028 0.062 0.182

0.019 0.019 0.019 0.019 0.019 0.019

−0.30 0.61 1.54 1.42 3.21 9.37

0.999 0.997 0.722 0.792 0.023 < 0.001

Table 12: RCA/SwiGLU symbol-retrieval contrasts from the BLiMP subtest binomial GLMM. Contrasts are relative to the relative (learned relative symbol) baseline; p-values are Tukey-adjusted.

Symbol condition

BLiMP

Entity

COMPS

relative (learned) relative_rope positional positional_sinusoidal symbolic relsymbolic relsymbolic_n4

69.34 69.43 69.17 68.88 68.90 68.33 66.32

22.13 20.54 16.69 23.12 17.70 21.17 19.88

52.56 52.54 52.56 52.14 52.60 52.09 51.63

Table 15: RCA/SwiGLU symbol-retrieval comparison, using 5 seeds on BLiMP, COMPS and Entity accuracy scores. All models use the same 9SA/3RA RCA architecture with RoPE positional encodings and SwiGLU activations. Further comparisons are in the appendix.

Symbol condition

Eye-tracking

SPR

relative (learned) relative_rope positional positional_sinusoidal symbolic relsymbolic relsymbolic_n4

9.94 9.41 9.78 9.84 9.64 9.93 10.06

3.64 3.66 3.81 3.70 3.58 3.85 3.90

Table 16: RCA/SwiGLU symbol-retrieval comparison (complete design: 7 symbol conditions × 5 seeds). All models use the same 9SA/3RA RCA architecture with RoPE positional encodings and SwiGLU activations. The relative condition is the reference baseline for contrasts. Shown measures are the eye-tracking and self-paced reading composite scores.

Effect (Intercept) symbol_retrieval

χ2

df

p

38.28 134.29

1 6

< 0.001 < 0.001

Table 17: Type III ANOVA (Wald χ2 tests) for the RCA/SwiGLU symbol-retrieval BLiMP subtest model (binomial GLMM). The symbol-retrieval mechanism has a large, significant effect on BLiMP subtest accuracy (χ2 (6) = 134.29, p < 0.001).

SA

RA

BLiMP

Supp.

EWoK

COMPS

1 3 4 6 8 9 11

11 9 8 6 4 3 1

68.22 68.22 69.23 70.33 69.18 69.70 69.67

56.92 57.38 56.92 56.32 57.06 57.75 56.87

50.01 50.37 50.44 50.55 50.75 50.60 51.13

51.67 51.70 51.77† 51.64 51.70 52.04 51.81

Table 18: RCA SA/RA head-ratio sweep at 512-token context, mean of 3 seeds. All models use RCA with SwiGLU and relative symbols. The 6SA/6RA split achieves the best BLiMP accuracy (70.33%). † One 4SA/8RA COMPS evaluation is missing; its COMPS mean is over two seeds.

Metric

Estimate (Slope) Std. Error

BLiMP Reading COMPS EWoK

-1.827 0.170 -0.245 -1.039

t

p

0.310 -5.887 <0.001*** 0.240 0.710 0.488 0.190 -1.287 0.201 0.447 -2.323 0.021*

Table 19: Mixed-effects models for the SA/RA headratio sweep at 512-token context (3 seeds per ratio, seed as random intercept). The RA head fraction has a significant negative effect on BLiMP and EWoK, but no significant effect on COMPS or reading measures. The seed random effect collapsed to zero variance in the BLiMP model.

Record · ID 978433 · SHA-256 ada3771fc5108444
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.