ConceptioArchivearXiv CS
arXiv CSopen access

MultiHashFormer: Hash-based Generative Language Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

MultiHashFormer: Hash-based Generative Language Models Huiyin Xue Atsuki Yamaguchi Nikolaos Aletras School of Computer Science, University of Sheffield, United Kingdom {hxue12, ayamaguchi1, n.aletras}@sheffield.ac.uk

arXiv:2606.28057v1 [cs.CL] 26 Jun 2026

Abstract Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To constrain the parameter footprint, prior work proposes hashing many tokens into a single vector within encoder-only models. While this offers parameter efficiency, many-to-one collisions prevent its use in causal LMs. In this paper, we propose M ULTI H ASH F ORMER, a new framework that allows hashbased autoregression. Each token is represented as a unique hash signature, a short sequence of discrete hash IDs, generated by multiple independent hash functions. A Hash Encoder compresses this signature into a single latent vector for processing by a Transformer decoder. Then, a Hash Decoder generates the hash signature of the next token, which is then mapped back to text. We evaluate our approach at the 100M, 1B and 3B parameter scales, demonstrating that M ULTI H ASH F ORMER consistently outperforms standard Transformer LMs across multiple benchmarks. Furthermore, we show that our model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.1

1

Introduction

Language models (Ettinger et al., 2025; Team, 2025; Abdin et al., 2024, LMs) typically map tokens into high-dimensional vectors via learned embedding matrices. This allows each token in a discrete vocabulary to be represented by a unique, dense vector. However, this linear scaling creates a vocabulary bottleneck, locking the model into a fixed token capacity and restricting its ability to adapt seamlessly to new domains, or languages. To mitigate this, previous research has explored token hashing (Prakash et al., 2020; Shu and Nakayama, 2018; Svenstrup et al., 2017; Ganchev and Dredze, 2008) to fix the parameter footprint of 1

MHF.

Code is available at https://github.com/HUIYINXUE/

the embedding space. Hash-based models such as the Proformer (Sankar et al., 2021) and the HashFormer (Xue and Aletras, 2022) use many-to-one mappings, where multiple tokens (e.g., cat, map, and physics) may share the same hash index (e.g., 40). While parameter-efficient, these models are restricted to encoder-based architectures and discriminative training. In generative settings, if a decoder predicts a shared hash index, the model cannot deterministically recover the intended token from the set of colliding candidates. Consequently, this ambiguity has made training causal LMs on hashed spaces practically infeasible. In this paper, we propose M ULTI H ASH F ORMER, a new framework designed to bypass the vocabulary bottleneck for decoder-based LMs. Drawing inspiration from chaotic dynamic memory systems (Skarda and Freeman, 1987) which leverage distributed, overlapping state spaces for highcapacity pattern retrieval (Bricken and Pehlevan, 2021), we replace the traditional embedding matrix with a modular hashing interface. Specifically, each token is converted into a unique hash signature consisting of a short sequence of discrete hash IDs generated by multiple independent hash functions, similar to chaotic transitions that allow biological networks to chain together vast vocabularies of concepts (Tsuda, 2001). For example, cat is represented as [12, 40, 56] and map as [99, 40, 3]. This multi-ID mapping eliminates the token collision problem of previous hash-based methods while allowing the model to scale its vocabulary capacity with a strictly sub-linear parameter footprint. For example, using four hash functions and 16,000 buckets per function, the M ULTI H ASH F ORMER theoretically supports an upper bound of 160004 (approx. 65 quadrillion) unique signatures. The interface is independent of the sequence processing backbone (e.g., a Transformer (Vaswani et al., 2017)) and manages the bidirectional translation between discrete tokens and the hash sig-

Figure 1: Overview of the M ULTI H ASH F ORMER framework using three different hash functions (for illustration purposes) with B buckets to obtain a multi-ID hash signature.

natures. At the input, a gated compositional embedding compresses the multi-ID signature into a dense latent vector, providing the backbone with a unified representation for sequence processing (Hash Encoder). At the output, a cascaded predictor reconstructs the hash signature of the next token sequentially, which is then deterministically mapped back to a specific discrete text token in the vocabulary (Hash Decoder). Figure 1 illustrates the M ULTI H ASH F ORMER architecture. Our main contributions are as follows: • We introduce the first hash-based framework that supports causal language modeling by preventing token collisions through multi-ID signature generation. • M ULTI H ASH F ORMER models consistently outperform standard Transformer LMs at 100M, 1B and 3B scales across 10 tasks, while offering better rare word representations. • We show that our models maintain performance while expanding the vocabulary size from 32K to 48K without any structural changes, or parameter count increase.

2

Related Work

2.1

Token-Free Models

LMs rely on an embedding matrix that scales linearly with the vocabulary size. While subword tokenizers like BPE (Sennrich et al., 2016) and SentencePiece (Kudo and Richardson, 2018) mitigate vocabulary explosion, they remain constrained by a fixed, data-derived vocabulary that struggles with rare or out-of-domain words. Token-free models such as AU-Net (Videau et al., 2025), Bolmo (Minixhofer et al., 2025), HNet (Hwang et al., 2025), BLT (Pagnoni et al., 2025), CANINE (Clark et al., 2022), and ByT5

(Xue et al., 2022) bypass the vocabulary bottleneck entirely by operating directly on unicode or byte sequences. However, this approach dramatically increases the input sequence length. Alternatively, T-FREE (Deiseroth et al., 2024) avoids tokenization by embedding words via sparse activations over locality-based hashed character trigrams. While parameter-efficient, it relies on characterlevel morphological similarity, making it languagedependent. In contrast, M ULTI H ASH F ORMER is orthogonal to these approaches. It resolves the vocabulary bottleneck while retaining the sequence compression advantages of subword tokenization. By decoupling the parameter matrix from the discrete vocabulary using combinatorial multi-hash mapping rather than sub-character heuristics, M UL TI H ASH F ORMER remains language-agnostic and supports any arbitrary tokenization strategy. 2.2

Factorized and Hash-Based Embeddings

To compress the memory footprint of standard embedding matrices, prior work has explored parameter reduction techniques. For example, ALBERT (Lan et al., 2020) uses matrix factorization to decouple the embedding dimension from the hidden dimension. While this reduces the parameter footprint, it still allocates a localized, explicit vector per token, failing to break the linear scaling constraint. A different approach is to use random hash embeddings (Svenstrup et al., 2017), which map discrete tokens into a highly compressed set of physical buckets. Architectural extensions like Proformer (Sankar et al., 2021) and HashFormer (Xue and Aletras, 2022) demonstrate that token hashing can be used to train Transformer-based models. However, compressing a vast vocabulary into a restricted physical bucket space inevitably forces multiple tokens to share the same index. This collision on a single hash function level prevents deterministic token recovery during

autoregressive generation, restricting these methods strictly to encoder-only architectures. M UL TI H ASH F ORMER resolves this limitation via the multi-identifier framework detailed in §3.

3

MultiHashFormer

M ULTI H ASH F ORMER comprises three modules shown in Figure 1: (1) a Hash Encoder that maps a discrete input token to a distributed multi-ID signature and compresses this signature into a single dense embedding; (2) a Sequence Processing Backbone that converts these embeddings into contextualized representations; and (3) a Hash Decoder that auto-regressively reconstructs the multiID signature of the next token. 3.1

Hash Encoder

Multi-Hash Indexing. To map each discrete token w into the hash signature space, we use H independent hash functions: H1 (w), H2 (w), . . . , HH (w). For any nonpadding token, the i-th coordinate of the signature is computed using the non-cryptographic MurmurHash3 (MMH3) algorithm (Senuma, 2025). The algorithm uses iterative bitwise multiplication and rotation to diffuse the input, ensuring a single-bit change yields a uniform, randomized hash. We use a hash function-specific seedi : Hi (w) = MMH3(w, seedi ) (mod B − 1) + 1

if w ̸= pad.

B denotes the number of discrete hash buckets per function. Hi (pad) = 0 is reserved across H to denote the padding token. To guarantee that every token in the vocabulary maps to a collision-free multi-hash ID, we use an iterative rehashing strategy. If a new token generates a signature that conflicts with an existing vocabulary entry, we incrementally modify the seed of the final hash function (seedH ) until an unused signature is found. Gated Compositional Embedding. We allocate H separate embedding matrices E(i) ∈ RB×d (where i ∈ {1, . . . , H}) to each hash coordinate H, where d represents the backbone hidden dimension. Because hash buckets are shared globally per function, overlapping tokens that are semantically unrelated inevitably collide. Inspired by the multi-embedding approach of Guo et al. (2024), we resolve this ambiguity by compressing the coordinates into a unified token representation

via a context-aware compositional gate. To determine the contribution of each bucket, the individual hash embeddings pass through a feed-forward bottleneck network with a compression dimension dz ≪ d, followed by softmax normalization. Finally, a linear adapter matrix Ws projects the combined representation into the latent space of the sequence processing backbone: e=

H X

! (i) α̃i E[Hi (w),:]

Ws ,

i=1

  expαi (i) , αi = σ E[Hi (w),:] W1 W2 . α̃i = PH αj j=1 exp where Ws ∈ Rd×d represents the structural adapter, W1 ∈ Rd×dz and W2 ∈ Rdz ×1 sequentially project the representations to an activation scaler with an non-linear activation above the intermediate bottleneck mainifold. Given an input sequence of tokens (w1 , w2 , . . . , wn ), the Hash Encoder processes each token wi to its corresponding hash embedding ei ∈ Rd . 3.2

Sequence Processing Backbone

X = [e1 , e2 , . . . , en ]⊤ where X ∈ Rn×d passes through a standard stack of L Transformer layers. At each time-step t, the final layer of the backbone emits a contextualized latent vector ht ∈ Rd : ht = Transformer(X[:t,:] ). Subsequently, ht is fed to the Hash Decoder. 3.3

Hash Decoder

To map the continuous hidden states back into the multi-hash ID signature space, we implement an auto-regressive Cascaded Predictor. This module iteratively refines its index predictions. It functions as a structured error-correcting system where preceding hashing choices directly constrain and contextualize subsequent token-signature inferences. To anchor this iterative cascade to the sequence context for predicting the next token wt+1 , the decoder initializes its prediction loop with the final hidden state of the backbone at the current timestep t. Thus, the root state c(1) ∈ Rd for the logit computation cascade is defined as c(1) = ht . Logit Computation. For each sequential hash head i ∈ {1, . . . , H}, the decoder projects the current latent hash state c(i) ∈ Rd to compute a logit distribution o(i) ∈ RB over the physical bucket allocations. This is executed by slicing a dedicated

(i)

subspace of the hash head weights Wo ∈ Rd×B : o(i) = Wo(i)⊤ c(i) , (i)

where Wo defines the localized parameter weight matrix assigned exclusively to the i-th coordinate bucket array. Following Press and Wolf (2017), we tie the input and output embedding weights by (i)⊤ setting Wo = E(i) . Soft Embedding Retrieval. For all non-terminal heads (i < H), the decoder extracts a continuous representation of its prediction to pass down the cascade. We first compute a normalized probability distribution p(i) ∈ RB across the bucket candidates via a softmax operation:

Training Phase. The Hash Decoder maps tokens into a virtual vocabulary space Vvirt defined by the total combinations of physical bucket allocations, where |Vvirt | = B H . Because the actual language vocabulary is highly compact relative to this space, it forms a strict subset: Vactl ⊂ Vvirt and |Vvirt | ≫ |Vactl |. Consequently, we allow predictions over coordinate combinations that might not map to actual entries in the language vocabulary. This unconstrained formulation simplifies the optimization objective. The training probability distribution for a given token w is thus optimized via the product of its independent coordinate probabilities: Ptrain (w | ·) =

p(i) = Softmax(o(i) ).

H Y

P (Hi (w) | ·).

i=1

Using a hard discrete index during intermediate steps disrupts differentiability. Instead, the module computes a soft bucket embedding e(i) ∈ Rd as the expected value of the shared embedding matrix E(i) , weighted by the logit probabilities: e(i) = E(i)⊤ p(i) State Update via Cascade Mixer. To propagate the contextualized trajectory to the next signature head, the internal hash state requires an update. We achieve this using a recursive cascade mixer inspired by tree-structured recursive networks (Socher et al., 2011; Goller and Küchler, 1996). This block concatenates the existing state with the retrieved soft embedding and routes the joint representation through a bottleneck layer. Augmented with a structural residual connection, the subsequent state formulation c(i+1) is defined as:  (i)  (i)⊤ c (i+1) (i) (i)⊤ c = c + Wup σ(Wdn ), e(i) where σ is an element-wise activation, and (i) (i) Wdn ∈ R2d×b and Wup ∈ Rb×h represent the low-rank projection weights of the bottleneck layer. 3.4

Probability Modeling

Standard autoregressive LMs decompose the joint probability of the input into a product of conditional probabilities. In contrast, M UL TI H ASH F ORMER leverages the outputs of the Hash Decoder to factorize this token-level prediction. Specifically, the normalized distributions (p(1) , . . . , p(H) ) produced iteratively by the cascaded predictor serve directly as the hash-specific conditional probabilities P (Hi (wt ) | ·) for each signature coordinate i ∈ {1, . . . , H}.

Inference Phase. Allowing unconstrained decoding at inference could cause the model to generate coordinate combinations that do not map to actual semantic entries, leaking probability mass into invalid signatures. To guarantee the generation of valid multi-hash ID sequences, we explicitly exclude all unassigned signature during inference. This requires re-normalizing the probability distribution strictly over the true token vocabulary Vactl : QH

i=1 P (Hi (w) | ·) QH ′ i=1 P (Hi (w ) | ·) w′ ∈Vactl

Pinf (w | ·) = P

By accumulating the log-probabilities of individual hash IDs, the final step is a standard softmax normalization over the valid token space Vactl .

4

Experimental Setup

4.1

Models

We evaluate M ULTI H ASH F ORMER models against standard autoregressive LMs across 100M, 1B and 3B parameter scales. All configurations use a decoder-only Transformer backbone based on the Qwen3 (Team, 2025) architecture. We use Mistral7B-v0.3 (Jiang et al., 2023), a predominantly English BPE tokenizer with a 32K vocabulary. Baselines. We use two baseline configurations: (1) A causal LM with a conventional embedding matrix E ∈ R|V|×d and an equivalent LM head Wo ∈ Rd×|V| (Standard); (2) a variant of Standard with k additional transformer layers to match the total parameter count of M ULTI H ASH F ORMER (Standard+kL). This allows us to verify that

any performance improvements in M ULTI H ASH F ORMER do not result from increased parameter capacity.2 Embedding and LM head weights are tied similar to the M ULTI H ASH F ORMER. We match the hidden dimension d and the number of attention heads to the M ULTI H ASH F ORMER. MultiHashFormer. We use the same core sequence processing backbone to the baseline configurations. We denote variants as HHBB, where H is the number of independent hash functions and B is the physical bucket allocation per head. We evaluate two configurations: H3B10K (H = 3, B = 10, 240)2 , which directly matches the baseline parameter count, and H4B16K (H = 4, B = 16, 384), our optimal configuration based on §6. For brevity, we fix the embedding matrices and projection layer sizes across all vocabulary settings. The bottleneck projection dimension for both the Gated Compositional Embedding and the Cascade Mixer is set to dz = 64 (where dz ≪ d). 4.2

Pre-training Data and Hyperparameters

All models are pre-trained from scratch on a subset of English FineWeb-Edu (Penedo et al., 2024). Following Chinchilla scaling laws (Hoffmann et al., 2022), we adjust data volume based on model scale. The 100M models were trained on 10B tokens, and the 1B and 3B models on 100B tokens. We use a global batch size of 256 with a 2,048-token context window. See Table 7 in Appendix C for details. 4.3

Evaluation Benchmarks

Core capabilities. We include LAMBADA (Paperno et al., 2016) for Language Modeling; Commonsense Reasoning (CR) using ARC-Easy (Bhakthavatsalam et al., 2021), COPA (Roemmele et al., 2011), OBQA (Mihaylov et al., 2018), PIQA (Bisk et al., 2020), HellaSwag (Zellers et al., 2019); and Reading Comprehension (RC) with RACE (Lai et al., 2017), SciQ (Welbl et al., 2017), SIQA (Sap et al., 2019). We evaluate CR + RC using ReCoRD (Zhang et al., 2018).3 Rare Words. To investigate if the shared bucket mechanism of M ULTI H ASH F ORMER improve rare word representations, we evaluate 1B models on 2 Due to computational resource constraints, we consider this variant for 100M and 1B scales only. 3 We report normalized accuracy for ARC-E, HellaSwag, SIQA, OBQA, and SciQ; F1 scores for ReCoRD; and standard accuracy for all remaining tasks. We use a five-shot setting for ARC-E, SIQA, and OBQA, while the rest are evaluated in a zero-shot setting.

the Card-660 dataset (Pilehvar et al., 2018). This dataset contains word pairs (rare-rare or rarefrequent) with human-annotated similarity scores. We compute the cosine similarity between the words using the final and second-to-last hidden states corresponding to the last token. Finally, we measure the correlation between these modelderived scores and the human annotations. 4.4

Vocabulary Expansion

As mentioned in §1, a key theoretical advantage of M ULTI H ASH F ORMER is its ability to expand its vocabulary capacity without increasing its parameter footprint. To evaluate this capability, we test the models on their ability to integrate new tokens during multilingual vocabulary expansion. Continual Pre-training Data. We continue pretraining (CPT) models on a multilingual corpus containing 6B tokens evenly split across Arabic, Chinese, and Hindi from FineWeb2 (Penedo et al., 2025). We also include an extra 2B English tokens from FineWeb-Edu (Penedo et al., 2024) to mitigate catastrophic forgetting. Following Yao et al. (2021), we add 5K tokens per language, yielding a final vocabulary of 48K tokens. Model Configuration. For M ULTI H ASH F ORMER, we expand the vocabulary size by registering new unique virtual (multi-hash ID) signatures without adding new parameters. For baselines, the weights of the new tokens within the embedding matrix are initialized using mean initialization, a standard protocol where each new token receives the average embedding of its corresponding source tokens from the original tokenizer (Mundra et al., 2024; Tejaswi et al., 2024; Yamaguchi et al., 2025, 2026). To reduce computational overhead, we only update the embedding layer, the LM heads for the Standard baselines, and the hash encoder and decoder for M ULTI H ASH F ORMER. Following Remy et al. (2024); Owodunni and Kumar (2025); Nakash et al. (2025), the first two and last two Transformer layers are also updated. See Table 8 (Appendix C) for full details. Evaluation. We use five multilingual tasks from MuBench (Han et al., 2025), which are HellaSwag, TruthfulQA (Lin et al., 2022), StoryCloze (Mostafazadeh et al., 2016), and

#Param

3B

1B

100M

Model

Emb

Lang. Model.

Dec

LAMBADA

Commonsense Reasoning ARC-E

COPA

OBQA

PIQA

HellaSwag

Reading Comprehension

CR+RC

RACE

ReCoRD

SciQ

SIQA

Rnd. Guess

-

-

-

25.00

50.00

25.00

50.00

25.00

25.00

25.00

25.00

19.10

Standard Standard+4L

25M 25M

87M 116M

15.33 19.64

41.33 45.08

64.00 64.00

27.00 27.00

58.87 60.17

28.59 29.99

26.41 26.79

59.00 62.60

35.98 36.80

49.47 53.08

MHF (H3B10K) MHF (H4B16K)

25M 52M

87M 87M

16.40 18.79

44.91 45.24

61.00 62.00

26.40 27.60

60.23 59.03

27.81 27.99

27.18 28.42

55.50 58.00

37.00 36.03

48.21 50.11

Standard Standard+2L

65M 65M

0.9B 1.0B

30.41 30.66

59.85 60.23

64.00 65.00

31.60 31.20

65.78 65.02

38.99 39.85

29.38 29.28

69.90 72.10

39.10 38.89

63.20 63.78

MHF (H3B10K) MHF (H4B16K)

65M 138M

0.9B 0.9B

35.78 35.34

61.95 64.35

65.00 71.00

34.00 29.80

68.17 66.49

40.45 41.32

29.00 31.39

68.30 70.20

40.48 41.66

64.51 64.90

Standard

65M

2.8B

28.64

60.23

65.00

31.80

65.83

39.65

29.76

64.80

38.28

62.26

MHF (H4B16K)

138M

2.8B

37.26

66.88

68.00

34.60

68.39

42.89

30.91

70.60

40.74

64.13

Table 1: Performance of standard and our M ULTI H ASH F ORMER models. Bold denotes best performance.

MMLU (Hendrycks et al., 2021).4 We also include the English benchmarks from §4.3 to track catastrophic forgetting.

5

Results

5.1

Core Capabilities

Table 1 presents the performance of the baselines alongside M ULTI H ASH F ORMER (MHF) models across the 100M, 1B, and 3B parameter scales. Efficacy while scaling parameter counts. We first observe that as the model capacity increases from 1B to 3B parameters, MHF (H4B16K) consistently outperforms the baseline on 9 out of 11 tasks. Specifically, on the LAMBADA task, MHF (H4B16K) offers substantial gains of 4.93% and 8.62% over the Standard baseline (30.41% and 28.64%) at the 1B and 3B scales, respectively. Although the standard LM baseline achieves higher scores on OBQA at the 1B scale, the overall results indicate the efficacy of M ULTI H ASH F ORMER in capturing contextual dependencies. Gains under strict parameter matching. When matching the parameter count at the 1B scale, M ULTI H ASH F ORMER variants offer superior performance compared to baselines. Specifically, MHF (H3B10K) and MHF (H4B16K) outperform Standard and Standard+2L baselines on 8 out of 10 tasks, respectively. Notably, MHF (H4B16K) achieves 64.90 on ReCoRD, surpassing the 63.78 of Standard+2L. Similarly, MHF (H3B10K) outperforms Standard on COPA with 71.00 vs. 65.00. These results suggest that decoupling the vocabulary is effective. Allocating more parameters to 4 We only include tasks where all Standard baselines outperform random guessing. For MMLU, we apply this filtering criterion at the subtask level and report the average over the retained subtasks. We report zero-shot accuracy for MMLU and normalized accuracy for the remaining tasks.

the embeddings improves performance more than adding more layers to the model on larger scales. Per-task performance. Looking at specific tasks, we note a performance trade-off based on the hash configuration. On OBQA, models with fewer hash buckets achieve higher accuracy. This reversal likely stems from the long-tail distribution of the dataset. A large bucket allocation fragments the embedding space for sparse vocabulary items, which degrades the representation of rare factual tokens. Conversely, reading comprehension requires high representational capacity to prevent coordinatelevel collisions between tokens across complex contexts. At the 1B scale, the lower-capacity MHF (H3B10K) variant underperforms the Standard baseline on SciQ (68.30% vs. 69.90%) and RACE (29.00% vs. 29.38%), although it outperforms the baseline on SIQA (40.48% vs. 39.10%). Expanding the physical bucket allocation to the MHF (H4B16K) configuration improves these scores to 70.20% on SciQ and 31.39% on RACE, surpassing the Standard baseline. M ULTI H ASH F ORMER’s strong performance on LM and reasoning tasks stems primarily from its ability to mitigate the softmax bottleneck (Yang et al., 2018), which is further analyzed in Appendix A. This advantage is especially evident in LAMBADA, which requires predicting specific, context-dependent words, and HellaSwag, which tests narrative logic. Our model effectively distinguishes between high-probability candidates that Standard baselines often conflate. Performance saturation at 100M. At the 100M scale, the decoder size is more important than the vocabulary representation. M ULTI H ASH F ORMER variants do not consistently outperform Standard baselines, especially on COPA, HellaSwag, and ReCoRD. Instead, increasing decoder depth by 4 layers is better across 7 tasks. We attribute this to

Figure 2: M ULTI H ASH F ORMER (blue shades, diamond marker) are similar or better than Standard (orange shades, circle marker) at 1B and 3B scales after expanding the vocabulary from 32K to 48K without adding a single new parameter. Black horizontal lines denote random baseline performance per task. Model

Last Hidden States Pearson r ↑ Spearman ρ ↑

Second-to-Last Hidden States Pearson r ↑ Spearman ρ ↑

Starndard MHF (H3B10K)

0.20 0.23

0.22 0.23

0.26 0.32

0.25 0.29

Standard+2L MHF (H4B16K)

0.22 0.20

0.22 0.26

0.29 0.30

0.29 0.32

Table 2: Correlations between human annotations and cosine similarities computed from the last two hidden states of the 1B models on the Card-660 dataset.

three constraints of smaller and shallower architectures: (1) restricted hidden dimensions conflate distinct tokens with similar hash signatures; (2) shallower depth prevents the model from contextualizing and disambiguating these embeddings; and (3) narrower hidden states restrict the rank upper bound (see our theoretical analysis in Appendix A), degrading log-probability estimation. Consequently, M ULTI H ASH F ORMER scales more effectively with larger decoders. 5.2

Rare Word Representations

Table 2 shows the correlation coefficients on the Card-660 dataset. The M ULTI H ASH F ORMER variants consistently outperform Standard models while maintaining an equivalent parameter count. We find that M ULTI H ASH F ORMER is better at identifying semantically equivalent word pairs (see examples in Appendix B). This performance gap is particularly evident when analyzing hidden states from the second-to-last decoder layer, where representations are less biased toward the specific training tasks. Our model effectively captures semantic representations of rare vocabulary items by forcing sparse tokens to share representational capacity. 5.3

Vocabulary Expansion

Figure 2 illustrates the performance of the adapted baselines compared to the adapted M ULTI H ASH -

F ORMER models at 1B and 3B parameter scales across Arabic, Hindi, Chinese, and English. Notably, both the 1B and 3B MHF (H4B16K) models consistently outperform Standard on 12 and 13 out of 22 multilingual tasks, respectively. They achieve this without the additional 31-million-parameters needed to accommodate 15K new tokens in the expanded vocabulary of the Standard baselines. We further find that MHF (H3B10K) and MHF (H4B16K) perform comparably (difference smaller than 1%) on 15 out of 22 tasks. This indicates that performance remains robust during vocabulary expansion. Furthermore, under a strict parametercount matching constraint prior to vocabulary expansion, MHF (H3B10K) and MHF (H4B16K) achieve better or comparable performance to the Standard models on the majority of English tasks (6 and 9 out of 10, respectively). This suggests that M ULTI H ASH F ORMER does not suffer from more catastrophic forgetting relative to Standard. It successfully supports a 46.9% larger multilingual vocabulary without a single additional parameter while preserving its core capabilities in English.

6

Analysis

We further analyze the core architectural components of M ULTI H ASH F ORMER. Due to computational resource constraints, we conduct this analysis at the 1B parameter trained up to 20B tokens. Single vs. Multi-hash ID. To demonstrate the necessity of Multi-ID signatures, we evaluate three different M ULTI H ASH F ORMER configurations (H4B4K, H4B8K, and H4B16K) against their Single-ID counterparts (H1B4K, H1B8K, and H1B16K). Table 3 presents the results on LAMBADA. Models with Multi-ID signatures consistently outperform those using Single-ID sig-

MHF

B4K

B8K

B16K

H1 H4

4.29 30.27

8.77 30.47

14.30 30.91

Table 3: LAMBADA accuracy across different MHF (H4BB) configurations against their corresponding variants H1BB using Single-ID signature.5

LSH : MMH3

0:4

1:3

2:2

3:1

4:0

LAMBADA

30.91

30.60

30.80

31.40

31.36

Table 4: LAMBADA accuracy across different MHF(H4B16K) variants while replacing a different number of MMH3 functions to LSH.

accuracy scores remain stable. Specifically, the best-performing variant improves over the lowest by only 4%, despite a tenfold increase in parameter count from 10M to 102M. Consequently, the H4B16K variant achieves an optimal balance between parameter efficiency and predictive accuracy. We therefore adopt H4B16K as the primary configuration for the 1B and 3B scaling experiments presented in §5.1. Figure 3: Average scores on four core capabilities across 1B M ULTI H ASH F ORMER variants at against their parameter counts in embeddings (vertical bars). Horizontal dotted lines represent the Standard baseline.

natures across all configurations (B4K, B8K, and B16K). Specifically, the performance difference is higher at lower bucket capacities (B4K), where H4B4K achieves 30.27% accuracy compared to the 4.29% of H1B4K. Expanding the capacity to B16K with Single-ID only improves accuracy to 14.30%. These results suggest that a signature collision severely degrades the core capability of the model. Employing Multi-ID signatures to prevent such collisions is more effective than increasing the number of hash buckets. Varying Hash Functions and Bucket Size. To evaluate the impact of hash configuration on parameter efficiency and downstream performance, we test various combinations of hash functions (H) and bucket sizes (B). Figure 3 illustrates average scores on four core capabilities of M ULTI H ASH F ORMER variants against their embedding parameter counts, across different combinations of H and B. All M ULTI H ASH F ORMER configurations consistently outperform Standard on Language Modeling and Commonsense Reasoning, including variants that utilize strictly smaller embedding matrices, such as H3B4K, H4B4K, and H3B8K. Although scaling both H and B sharply increases the embedding parameter count, the corresponding 5

H1BB use single value for 4 coordinates to strictly match the number of parameters in H4B for a fair comparison.

MMH3 vs. LSH. We also investigate whether incorporating Locality-Sensitive Hashing (Datar et al., 2004, LSH) enhances the general capability of the model by capturing morphological similarity. Because LSH maps structurally similar input features to identical hash buckets with high probability, applying it to character-level trigrams tends to force subwords with similar spelling to share embedding representations. To evaluate this mechanism, we pre-train three additional MHF (H4B16K) variants from scratch at a 1B-parameter scale, keeping the dataset and hyperparameters identical. Table 4 presents the performance of these models on LAMBADA. Increasing the proportion of LSH functions while fixing the total number of hash functions does not yield a consistent accuracy improvement. This outcome suggests that explicitly injecting a morphological inductive bias is not necessary given the inherent computational complexity of LSH compared to MMH3. The standard deterministic random hashing of MMH3 provides sufficient representational flexibility for the network to learn subword relationships entirely end-to-end.

7

Conclusion

We introduced M ULTI H ASH F ORMER, a generative framework that leverages multi-hash structures to bypass traditional vocabulary bottlenecks. Empirical evaluation demonstrates that it offers improvements both in core capabilities and rare word representations. Furthermore, the framework enables seamless vocabulary expansion without requiring additional parameters.

Limitations Model size. Due to computational resource constraints, we restrict our evaluation to models at the 100M, 1B and 3B scales. Following the protocol of Cheng et al. (2024), 1B and 3B models were trained from scratch within a standard academic budget of 100B tokens. While M ULTI H ASH F ROMER shows an upward performance trajectory with model size, exploring larger scales (e.g., 7B+ parameters) remains a valuable direction for institutions with access to the required computing resources. Single seed. Due to the high computational cost of pre-training 1B and 3B parameter models from scratch for 100B tokens, conducting multiple training runs was prohibitive under our academic infrastructure constraints. Consequently, our empirical results are based on a single seed per configuration. However, the consistent performance gains observed across both the 1B and 3B scales demonstrate the robustness of the M ULTI H ASH F ORMER framework. This consistency indicates that the improvements introduced by our Hash Encoder and Hash Decoder stem from inductive bias rather than random sampling variance. Furthermore, adopting the established hyperparameter protocols of Li et al. (2025) and Cheng et al. (2024), we ensure standard and reproducible training dynamics.

Acknowledgments We would like to thank Maggie Mi for the valuable feedback. We acknowledge (1) IT Services at the University of Sheffield for the provision of services for high-performance computing; (2) the use of the University of Oxford Advanced Research Computing (ARC) facility; (3) the Isambard-AI National AI Research Resource (AIRR), which is operated by the University of Bristol and is funded by the UK Government’s Department for Science, Innovation and Technology (DSIT) via UK Research and Innovation; and the Science and Technology Facilities Council [ST/AIRR/I-A-I/1023]; and (4) the EuroHPC Joint Undertaking for awarding us access to Leonardo, hosted by CINECA (Italy).

Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, and 8 others. 2024. Phi-4 technical report. CoRR, abs/2412.08905. Sumithra Bhakthavatsalam, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, and Peter Clark. 2021. Think you have solved direct-answer question answering? try arc-da, the direct-answer AI2 reasoning challenge. CoRR, abs/2102.03315. Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 7432– 7439. AAAI Press. Trenton Bricken and Cengiz Pehlevan. 2021. Attention approximates sparse distributed memory. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 15301–15315. Daixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi, Minlie Huang, and Furu Wei. 2024. Instruction pretraining: Language models are supervised multitask learners. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 2529–2550, Miami, Florida, USA. Association for Computational Linguistics. Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022. Canine: Pre-training an efficient tokenization-free encoder for language representation. Transactions of the Association for Computational Linguistics, 10:73–91. Mayur Datar, Nicole Immorlica, Piotr Indyk, and Vahab S. Mirrokni. 2004. Locality-sensitive hashing scheme based on p-stable distributions. In Proceedings of the 20th ACM Symposium on Computational Geometry, Brooklyn, New York, USA, June 8-11, 2004, pages 253–262. ACM.

References

Björn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting, and Samuel Weinbach. 2024. TFREE: Subword tokenizer-free generative LLMs via sparse representations for memory-efficient embeddings. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21829–21851, Miami, Florida, USA. Association for Computational Linguistics.

Marah I Abdin, Jyoti Aneja, Harkirat S. Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li,

Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini,

Matt Jordan, Mayee F. Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, and 48 others. 2025. Olmo 3. CoRR, abs/2512.13961. Kuzman Ganchev and Mark Dredze. 2008. Small statistical models by random feature mixing. In Proceedings of the ACL-08: HLT Workshop on Mobile Language Processing, pages 19–20, Columbus, Ohio. Association for Computational Linguistics. Christoph Goller and Andreas Küchler. 1996. Learning task-dependent distributed representations by backpropagation through structure. In Proceedings of International Conference on Neural Networks (ICNN’96), Washington, DC, USA, June 3-6, 1996, pages 347–352. IEEE. Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. 2024. On the embedding collapse when scaling up recommendation models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, Proceedings of Machine Learning Research, pages 16891–16909. PMLR / OpenReview.net. Wenhan Han, Yifan Zhang, Zhixun Chen, Binbin Li, Haobin Lin, Bingni Zhang, Taifeng Wang, Mykola Pechenizkiy, Meng Fang, and Yin Zheng. 2025. Mubench: Assessment of multilingual capabilities of large language models across 61 languages. CoRR, abs/2506.19468. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, and 3 others. 2022. An empirical analysis of compute-optimal large language model training. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022. Sukjun Hwang, Brandon Wang, and Albert Gu. 2025. Dynamic chunking for end-to-end hierarchical sequence modeling. CoRR, abs/2507.07955. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. CoRR, abs/2310.06825.

Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics. Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. RACE: Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785– 794, Copenhagen, Denmark. Association for Computational Linguistics. Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net. Houyi Li, Wenzhen Zheng, Jingcheng Hu, Qiufeng Wang, Hanshan Zhang, Zili Wang, Shijie Xuyang, Yuantao Fan, Shuigeng Zhou, Xiangyu Zhang, and Daxin Jiang. 2025. Predictable scale: Part I - optimal hyperparameter scaling law in large language model pretraining. CoRR, abs/2503.04715. Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics. Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium. Association for Computational Linguistics. Benjamin Minixhofer, Tyler Murray, Tomasz Limisiewicz, Anna Korhonen, Luke Zettlemoyer, Noah A. Smith, Edoardo M. Ponti, Luca Soldaini, and Valentin Hofmann. 2025. Bolmo: Byteifying the next generation of language models. CoRR, abs/2512.15586. Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, San Diego, California. Association for Computational Linguistics. Nandini Mundra, Aditya Nanda Kishore Khandavally, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan,

and Mitesh M Khapra. 2024. An empirical comparison of vocabulary expansion and initialization approaches for language models. In Proceedings of the 28th Conference on Computational Natural Language Learning, pages 84–104, Miami, FL, USA. Association for Computational Linguistics. Itay Nakash, Nitay Calderon, Eyal Ben-David, Elad Hoffer, and Roi Reichart. 2025. Adaptivocab: Enhancing LLM efficiency in focused domains through lightweight vocabulary adaptation. CoRR, abs/2503.19693. Abraham Toluwase Owodunni and Sachin Kumar. 2025. Continually adding new languages to multilingual language models. CoRR, abs/2509.11414. Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srini Iyer. 2025. Byte latent transformer: Patches scale better than tokens. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9238–9258, Vienna, Austria. Association for Computational Linguistics. Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1525–1534, Berlin, Germany. Association for Computational Linguistics. Guilherme Penedo, Hynek Kydlícek, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin A. Raffel, Leandro von Werra, and Thomas Wolf. 2024. The fineweb datasets: Decanting the web for the finest text data at scale. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. Guilherme Penedo, Hynek Kydlícek, Vinko Sabolcec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro von Werra, and Thomas Wolf. 2025. Fineweb2: One pipeline to scale them all - adapting pretraining data processing to every language. CoRR, abs/2506.20920. Mohammad Taher Pilehvar, Dimitri Kartsaklis, Victor Prokhorov, and Nigel Collier. 2018. Card-660: Cambridge rare word dataset - a reliable benchmark for infrequent word representation models. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1391–1401, Brussels, Belgium. Association for Computational Linguistics.

Prafull Prakash, Saurabh Kumar Shashidhar, Wenlong Zhao, Subendhu Rongali, Haidar Khan, and Michael Kayser. 2020. Compressing transformer-based semantic parsing models using compositional code embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4711– 4717, Online. Association for Computational Linguistics. Ofir Press and Lior Wolf. 2017. Using the output embedding to improve language models. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 157–163, Valencia, Spain. Association for Computational Linguistics. François Remy, Pieter Delobelle, Hayastan Avetisyan, Alfiya Khabibullina, Miryam de Lhoneux, and Thomas Demeester. 2024. Trans-tokenization and cross-lingual vocabulary transfers: Language adaptation of llms for low-resource NLP. CoRR, abs/2408.04303. Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In Logical Formalizations of Commonsense Reasoning, Papers from the 2011 AAAI Spring Symposium, Technical Report SS-11-06, Stanford, California, USA, March 21-23, 2011. AAAI. Chinnadhurai Sankar, Sujith Ravi, and Zornitsa Kozareva. 2021. ProFormer: Towards on-device LSH projection based transformers. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2823–2828, Online. Association for Computational Linguistics. Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social IQa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4463– 4473, Hong Kong, China. Association for Computational Linguistics. Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics. Hajime Senuma. 2025. mmh3: A python extension for murmurhash3. J. Open Source Softw., 10(105):6124. Raphael Shu and Hideki Nakayama. 2018. Compressing word embeddings via deep compositional code learning. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.

Christine A. Skarda and Walter J. Freeman. 1987. How brains make chaos in order to make sense of the world. Behavioral and Brain Sciences, 10(2):161– 173. Richard Socher, Cliff Chiung-Yu Lin, Andrew Y. Ng, and Christopher D. Manning. 2011. Parsing natural scenes and natural language with recursive neural networks. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28 - July 2, 2011, pages 129–136. Omnipress. Dan Svenstrup, Jonas Meinertz Hansen, and Ole Winther. 2017. Hash embeddings for efficient word representations. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 4928–4936. Qwen Team. 2025. Qwen3 technical report. CoRR, abs/2505.09388. Atula Tejaswi, Nilesh Gupta, and Eunsol Choi. 2024. Exploring design choices for building languagespecific LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 10485–10500, Miami, Florida, USA. Association for Computational Linguistics. Ichiro Tsuda. 2001. Toward an interpretation of dynamic neural activity in terms of chaotic dynamical systems. Behavioral and Brain Sciences, 24(5):793–810. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008. Mathurin Videau, Badr Youbi Idrissi, Alessandro Ferreira Leite, Marc Schoenauer, Olivier Teytaud, and David Lopez-Paz. 2025. From bytes to ideas: Language modeling with autoregressive u-nets. CoRR, abs/2506.14761. Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy Usergenerated Text, pages 94–106, Copenhagen, Denmark. Association for Computational Linguistics. Huiyin Xue and Nikolaos Aletras. 2022. HashFormers: Towards vocabulary-independent pre-trained transformers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7862–7874, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Linting Xue, Aditya Barua, Noah Constant, Rami AlRfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022. ByT5: Towards a token-free

future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10:291–306. Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, and Nikolaos Aletras. 2025. Adapting chat language models using only target unlabeled language data. Trans. Mach. Learn. Res., 2025. Atsuki Yamaguchi, Aline Villavicencio, and Nikolaos Aletras. 2026. How can we effectively expand the vocabulary of LLMs with 0.01GB of target language text? Computational Linguistics, 52(1):295–330. Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018. Breaking the softmax bottleneck: A high-rank RNN language model. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 May 3, 2018, Conference Track Proceedings. OpenReview.net. Yunzhi Yao, Shaohan Huang, Wenhui Wang, Li Dong, and Furu Wei. 2021. Adapt-and-distill: Developing small, fast and effective pretrained language models for domains. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 460–470, Online. Association for Computational Linguistics. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy. Association for Computational Linguistics. Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. Record: Bridging the gap between human and machine commonsense reading comprehension. CoRR, abs/1810.12885.

A

Softmax Bottleneck

A.1

Preliminaries

Lemma 1. Given a model parameter θ, Hθ Wθ⊤ ∈ F (A) if and only if Pθ (X|c) holds for all c ∈ L.

Language models (LMs) define the conditional distribution Pθ (w|c) by applying a softmax function to a linear projection of the hidden state ht : exp h⊤ t ww ⊤ ′ w′ exp ht ww

Pθ (w|c) = P

where both the context vector ht (ct ; θ) and the word embedding wwt (wt ; θ) are on a ddimensional space. Their inner product, h⊤ t wwt , defines the logit. To analyze the expressive capacity of the softmax function, we define the following three matrices and their log-probabilities: Hθ = [h1 , . . . , hn ]⊤ ,

Wθ = [w1 , . . . , w|V| ]⊤

A = [ai,j ]n×|V| ∈ Rn×|V| , where ai,j = log p∗ (wj |ci )  exp(h⊤ t ww ) log p (w|c) = log P ⊤ ′ w′ exp(ht ww ) X ⊤ = ht ww − log exp(h⊤ t ww ′ ) ∗



w′

Hθ ∈ Rn×d denote the matrix of ht and Wθ ∈ R|V|×d is the projection matrix for all tokens in the vocabulary. The matrix A ∈ Rn×|V| contains the log probabilities of the true data distribution. We index H and W by θ to indicate that both represent functions within a joint family U parameterized by θ. In practice, Hθ is realized via a Transformerbased LM, while Wθ is instantiated as a learned token embedding lookup table. We further define a set of matrices, F (A), generated by applying a row-wise shift to A: F (A) = {A + ΛJn,|V| | Λ ∈ Rn×n is a diagonal matrix}

where Jn,|V| ∈ Rn×|V| denotes the all-ones matrix. This row-wise shift operation effectively adds a unique scalar constant to each element of A’s rows. Since the diagonal entries of Λ can be any real values, F (A) constitutes an infinite set characterized by two key properties: Property 1. For any matrix A′ , A′ ∈ F (A) if and only if Softmax(A′ ) = P ∗ . Thus, F (A) represents the set of all logit matrices that recover the true data distribution. Property 2. For any A1 , A2 ∈ F (A) such that A1 ̸= A2 , it holds that |rank(A1 ) − rank(A2 )| ≤ 1. Hence, the matrices in F (A) maintain consistent ranks, with a maximum discrepancy of one. Given Property A.1, we establish:

Following this lemma, expressiveness is framed as follows: does there exist a parameter θ and a matrix A′ ∈ F (A) satisfying Hθ Wθ⊤ = A′ . LMs learn matrices Hθ and Wθ to factorize a target matrix A′ ∈ F (A). For a valid factorization, the rank of the product Hθ Wθ⊤ must be at least equal to the rank of A′ . Because Hθ ∈ Rn×d and Wθ ∈ R|V|×d , the rank of this product is strictly upper-bounded by the embedding dimension d. Consequently, if d ≥ rank(A′ ), a universal approximator can recover A′ . Conversely, if d < rank(A′ ), no pair (Hθ , Wθ ) can recover A′ , regardless of the expressiveness of U. Proposition 1. Assuming the function family U is a universal approximator, there exists a parameter θ such that Pθ (X|c) = P ∗ (X|c) for all c ∈ L if and only if d ≥ minA′ ∈F (A) rank(A′ ). By combining Proposition A.1 with the properties of F (A) outlined in Property A.1, Yang et al. (2018) define the Softmax Bottleneck: Corollary 1. (Softmax Bottleneck). If d < rank(A) − 1, then for any function family U that acts as a universal approximator, there is no parameter θ such that Pθ (X|c) = P ∗ (X|c) for all c ∈ L. That is, Pθ (X|c) ̸= P ∗ (X|c) must hold for at least some contexts. This demonstrates that an insufficient dimension d restricts Softmax from expressing the true data distribution. The conclusion is not restricted to a finite language L. Because when L is infinite, one can always take a finite subset and the softmax bottleneck still exists. A.2

M ULTI H ASH F ORMER as a Mitigation Strategy While training, M ULTI H ASH F ORMER approximates elements in A by log p∗ (w|c) = log

H Y p∗ (Hi (w)|c) i

=

H X

log p∗ (Hi (w)|c)

i=1

=

H X

(i)

log

i=1

=

(i)

exp(c(i)⊤ wHi (w) )πHi (w) PB (i)⊤ w(i) ) j=1 exp(c

H X (i) (i) c(i)⊤ wHi (w) πHi (w) i=1

H X i=1

log

B X j=1

(i)

exp(c(i)⊤ wj )

!

The model projects representations to coordinates of the hash signature, using H localized parameter weight matrices assigned exclusively to i-th coor(i) dinate bucket array, Wo ∈ Rd×B . These coordinates are then mapped to the true vocabulary space via transition matrices π (i) ∈ {0, 1}B×|V| . Here, each column of π (i) is a one-hot vector, satisfying PB (i) j=1 πj,k = 1 for every column k. The estimation could be further written in a matrix form: ÂM HF =

H X

H X i=1

log

B X

! (i)

exp(Hθ Wo(i)⊤ )

j=1

Consequently, rank(ÂMHF ) ≤ min(n, B, |Vactl. |, H × min(d, B)) = min(B, H × d), where H denotes the number of hash functions configured by M ULTI H ASH F ORMER, B denotes the number of discrete hash buckets per function, and d is the dimensionality of the latent hash state vectors c(i) . Given that the rank of Standard baselines is bounded by rank(ÂStandard ) ≤ min(d, n, |Vactl. |) = d, M ULTI H ASH F ORMER elevates the upper bound of this rank to enhance model expressiveness, while employing multi hash functions and non-linear transformation between latent hash state vectors. This aligns with the intuition that aggregating multiple distinct observations inherently improves estimation accuracy. During inference, although we exclude all unassigned signatures, the rank upper bound remains invariant between training and inference. This stability stems from the fact that the row-wise partition function does not contribute to the final upper bound of the rank. While MultiHashFormer does not produce an arbitrary rank, it elevates the rank of the conditional distribution matrix. This improvement drives performance gains in highentropy reasoning tasks, including LAMBADA and HellaSwag.

B

Human

Stnd.

H3B10K

Stnd.+2L

H4B16K

Abbrevation retweeting ↔ RTing science-fiction ↔ sci-fi Brooklyn ↔ Brklyn Internet Explorer 4 ↔ IE4 ITV2 ↔ ITV Two electrolytic polishing ↔ electropolishing

1.00

0.56

0.65

0.47

0.71

1.00

0.77

0.90

0.77

0.86

1.00

0.45

0.53

0.46

0.61

1.00

0.51

0.61

0.53

0.61

1.00

0.51

0.60

0.43

0.52

0.99

0.74

0.83

0.62

0.87

Alias (i) Hθ Wo(i)⊤ π (i)

i=1

Word Pair

Rare Word Examples from Card-660

We first look to word pairs with high humanannotated similarity to serve as hard examples for our illustration. Table 5 presents examples of highsimilarity word pairs from the Card-660 dataset. We compare normalized human annotation scores against cosine similarities derived from the secondto-last decoder layers of the 1B Standard baseline and M ULTI H ASH F ORMER. Compared to the

full-HD ↔ 1080p Malva parviflora ↔ cheeseweed first milk ↔ colostrum little bee-eater ↔ Merops pusillus

1.00

0.55

0.72

0.40

0.67

1.00

0.52

0.58

0.57

0.61

1.00

0.54

0.64

0.53

0.69

1.00

0.48

0.60

0.51

0.64

Misspelling leggin ↔ legging tariqa ↔ tariqah shit ↔ shxt sweeeet ↔ sweet

1.00

0.59

0.57

0.39

0.60

1.00

0.77

0.82

0.72

0.82

0.99

0.17

0.53

0.22

0.52

0.99

0.36

0.45

0.37

0.52

Synonym 25 ↔ twenty-five shapeless ↔ amorphous unforeseen ↔ unanticipated

1.00

0.75

0.81

0.75

0.84

0.99

0.38

0.50

0.51

0.63

0.97

0.82

0.91

0.83

0.93

Table 5: Examples of word pairs from Card-660 with high human-annotated similarity scores (normalized to [0, 1]), compared against cosine similarities computed from the second-to-last decoder layer of the 1B Standard baselines and M ULTI H ASH F ORMER models.

Standard baseline, M ULTI H ASH F ORMER consistently assigns higher similarity scores to semantically equivalent word pairs, such as abbreviations, aliases, misspellings, and synonyms. Examples from word pairs with low similarities and medium similarities are shown in Table 6.

Word Pair

Human

Stnd.

H3B10K

Stnd.+2L

H4B16K

C.2

Low NetMeeting ↔ Marwar Hall Park Ji-sung ↔ Yosemite Park Pizza Hut ↔ Pizzle rot cheddah ↔ cheddar cheddah ↔ cheddar Apple ↔ Applebees

0.00

0.37

0.41

0.39

0.34

0.00

0.36

0.49

0.36

0.43

0.02

0.29

0.25

0.31

0.27

0.06

0.44

0.23

0.47

0.51

0.06

0.44

0.23

0.47

0.51

0.10

0.49

0.56

0.49

0.65

Medium Ben-Hur ↔ Titanic radionavigation ↔ frequency band night sky ↔ skyglow circus ↔ ropedancer transmigration ↔ residence permit Head tilt ↔ cervix

0.39

0.40

0.39

0.27

0.50

0.44

0.37

0.44

0.34

0.42

0.49

0.60

0.65

0.60

0.64

0.50

0.46

0.56

0.51

0.47

0.52

0.44

0.64

0.50

0.57

0.52

0.35

0.49

0.34

0.49

Table 6: Examples of word pairs from Card-660 with low and medium human-annotated similarity scores (normalized to [0, 1]), compared against cosine similarities computed from the second-to-last decoder layer of the 1B Standard baselines and M ULTI H ASH F ORMER models. A human annotation score of 1.00 indicates semantic equivalence.

Hyperparameters for VE Continual-pretraining

Hyperparameters

1B

Adaptive Decoder Layer Indices Attention dropout Tie word embeddings Vocabulary size Tokenizer

[1,2,19,20] [1,2,35,36] 0.0 0.0 True True 48,122 48,122 expanded expanded Mistral Mistral 256 256 16K 16K 2,048 2,048 3e-4 2e-4 cosine cosine 800 800 AdamW AdamW 1e-8 1e-8 0.9 0.9 0.999 0.999 1.0 1.0 0.1 0.1 BF16 BF16

Batch size Train steps Sequence length Maximum Learning Rate Learning rate scheduler Warmup steps Optimizer Adam ϵ Adam β1 Adam β2 Gradient clipping Weight decay Training precision

3B

Table 8: Additional hyperparameters of continualpretraining at each model scale.

D

MultiHashFormer Abbrevations and Configurations

Table 9 presents the detailed configurations of M ULTI H ASH F ORMER variants along with their abbreviations.

C

Hyperparameters

C.1

Hyperparameters for Pre-training

Hyperparameters

100M

1B

3B

Hidden size Intermediate size Max window layers Added layers (if scaling the depth) Number of attention heads Number of hidden layers Number of key value heads Rope theta RMS norm eps Attention dropout Tie word embeddings Hidden activation Initializer range Vocabulary size Tokenizer Batch size Train steps Sequence length Maximum Learning Rate Learning rate scheduler Warmup steps Optimizer Adam ϵ Adam β1 Adam β2 Gradient clipping Weight decay Training precision

768 2560 12 4 12 12 2 1M 1e-06 0.0 True SiLU 0.02 32,768 Mistral 256 20K 2,048 5e-4 cosine 2000 AdamW 1e-8 0.9 0.999 1.0 0.1 BF16

2048 6144 20 2 16 20 8 1M 1e-06 0.0 True SiLU 0.02 32,768 Mistral 256 200K 2,048 3e-4 cosine 2000 AdamW 1e-8 0.9 0.999 1.0 0.1 BF16

2048 11008 36 1 16 36 2 1M 1e-06 0.0 True SiLU 0.02 32,768 Mistral 256 200K 2,048 2e-4 cosine 2000 AdamW 1e-8 0.9 0.999 1.0 0.1 BF16

Table 7: Hyperparameters of pre-training at each model scale.

MHF abbr.

H

B

H3B4K H3B8K H3B10K H3B16K H4B4K H4B8K H4B16K H4B32K

3 3 3 3 4 4 4 4

4,096 8,192 10,624 16,384 4,096 8,192 16,384 32,768

Table 9: Detailed configurations of M ULTI H ASH F ORMER variants along with their abbreviations.

Record · ID 319685 · SHA-256 38b8f083fee14822
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.