Complexity-Guided Component-wise Initialization for Language Model Pretraining Konstantin Garbers∗
Nicholas Oh∗
Peking University Beijing, China [email protected]
Peking University Beijing, China [email protected]
arXiv:2607.09204v1 [cs.CL] 10 Jul 2026
Abstract Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-style language-model pretraining. First, we analyze eleven pretrained GPT-2-style checkpoints that vary in size, language, tokenizer, and training corpus, measuring Frobenius norm and effective-rank entropy across layers and Transformer subcomponents. The checkpoints show shared depth trends, especially increasing scale and stronger spectral concentration in residualwriting matrices. We then construct initialization schemes that imitate the component-wise magnitudes and spectral profiles of pretrained models, and compare them with several weight initialization methods. These initializers visibly change the model’s structural spectral patterns, but the evaluation results do not show a corresponding performance advantage. Pretrained-weight reuse remains competitive, while coarse spectral matching alone is not a reliable optimization strategy. Our results suggest that pretrained spectra are useful diagnostics of trained model structure, but that effective reuse likely requires preserving richer information than component-wise scale and singular-value shape.
of trained transformers. At the same time, it is unclear whether such coarse spectral information is actually useful for pre-training. Spectral similarity may capture meaningful structure, but it may also discard the specific directions, correlations, and feature-level organization that make pre-trained weights effective. In this paper, we study this question for GPT-2-style language models. We analyze the weight matrices of several pre-trained models using spectral tools, including the Frobenius norm and effective rank entropy [16]. Our goal is first to determine whether consistent spectral patterns appear across models that differ in size, language, tokenizer, and training corpus. We then test whether these patterns can be transferred into the initialization of a new model and whether such initialization improves pre-training dynamics. Research questions and answers. Specifically, we ask the following questions. Research question
Answer
RQ1
What spectral patterns appear in pre-trained GPT-2style language models?
RQ2
How does initialization based on these spectral patterns affect training dynamics?
Pre-trained GPT-2-style models show recurring layerwise and component-wise spectral trends despite differences in size, language, tokenizer, and corpus. Reusing pre-trained weights remains competitive, but copying coarse spectral shape or scale alone is not a reliable optimization strategy.
Keywords language model pretraining, transformer initialization, layerwise diagnostics, memorization, training dynamics
1
Introduction
Large language models are typically trained from randomly initialized weights, often drawn from Gaussian or related distributions. This choice is robust and architecture-agnostic: it does not assume prior knowledge about the data, the model size, or the internal structure that training will eventually produce. However, it also ignores a growing body of evidence suggesting that trained language models are not arbitrary points in parameter space. Across models, layers, and components, pre-trained transformers often exhibit recurring structural regularities, including characteristic spectral properties of their weight matrices [16, 29, 33]. These regularities raise a natural question: if trained models repeatedly develop similar patterns, can some of these patterns be used before training begins? In principle, an initialization that already reflects common structure found in pre-trained models could reduce the burden on optimization. Instead of learning all structural properties from scratch, the model would start from weights whose scale, rank structure, or spectral shape better resembles those ∗ Both authors contributed equally to this research.
The remainder of this paper is organized as follows. We first analyze existing pre-trained LLMs in section 2 to identify recurring spectral patterns in their weights. We then propose and evaluate a spectral-pattern-based initialization method in section 3. Finally, we discuss related work in section 5 and conclude in section 6.
2
Spectral Patterns in Pre-trained LLMs
Spectral analysis is the analysis of the eigenvalues or singular values of a matrix. In the context of neural networks, spectral analysis can provide insight into the geometric and functional properties of a weight matrix. Large singular values indicate that the matrix strongly amplifies certain input directions, which is often associated with greater complexity or sensitivity [6, 16, 25]. We define spectral patterns in section 2.2 and analyze pretrained GPT-2-style models in section 2.3 to identify shared spectral patterns in pre-trained LLMs.
Konstantin Garbers and Nicholas Oh
Component 𝑊𝑄 ,𝑊𝐾 𝑊𝑉 ,𝑊𝑂
𝑊up 𝑊down
Association used in interpretation
2.2
Larger singular values are associated with stronger feature directions for attention addressing [3, 34]. Larger singular values are associated with stronger feature directions for transported content and stronger write strength to the residual stream [2, 3, 25]. Component weights are usually associated with MLP key weights or feature-detector weights [4, 5]. Component weights are usually associated with MLP value weights or stored-feature readout weights [4, 5].
We use static spectral metrics because they are inexpensive to compute directly from saved weight matrices and do not require additional data passes, interventions, or fine-tuning runs. They are also established diagnostics for trained neural networks and Transformer weights: prior work uses spectral summaries to study implicit regularization, layer and component structure, initialization scale, and complexity control [16, 25, 29, 33, 35, 36]. Our chosen metrics separate complementary properties of each matrix: Frobenius norm captures total scale, while effective-rank entropy describes how concentrated or distributed that scale is across singular directions.
Table 1: Interpretive associations for component weights and singular values in Transformer block components.1
Spectral Analysis Metrics
Frobenius Norm. The Frobenius norm is √︄∑︁ √︄∑︁ 𝑊𝑖2𝑗 = ∥𝑊 ∥ 𝐹 = 𝜎𝑘 (𝑊 ) 2 . 𝑖,𝑗
𝑥˜ ℓ
𝑊𝑄𝐾𝑉 𝑥ℓ
LN1
Attn
+
LN2
𝑊𝑂
𝑊up 𝜙 𝑊down
+
𝑥 ℓ +1
Figure 1: Pre-layer-norm GPT-2 block notation. Attention and MLP outputs are added back into the residual stream.
2.1
GPT-2 Block Notation
𝑘
It measures total weight energy and is useful for separating changes in scale from changes in spectral shape. Effective Rank Entropy. Frobenius norm alone does not capture the distribution of singular values. Let 𝜎𝑖 be the singular values of Í 𝑊 , and let 𝑝𝑖 = 𝜎𝑖 / 𝑗 𝜎 𝑗 . We use the entropy effective rank ! ∑︁ erank(𝑊 ) = exp − 𝑝𝑖 log 𝑝𝑖 . 𝑖
GPT-2 is a decoder-only Transformer: tokens are embedded into a residual stream, then passed through a stack of identical causal self-attention and MLP blocks before the final language-model head. Each block first applies layer normalization and multi-head causal self-attention, writes the attention output back to the residual stream, then applies a second layer normalization and a positionwise MLP with another residual write. For block ℓ, let 𝑥 ℓ denote the residual stream entering the block. Using the pre-layer-norm GPT-2 convention, the two residual writes are 𝑎 ℓ = Attn(LN1 (𝑥 ℓ );𝑊𝑄𝐾𝑉 ,𝑊𝑂 ),
𝑥˜ℓ = 𝑥 ℓ + 𝑎 ℓ ,
𝑚 ℓ = 𝑊down𝜙 (𝑊up LN2 (𝑥˜ℓ )),
𝑥 ℓ+1 = 𝑥˜ℓ + 𝑚 ℓ ,
where 𝜙 is the MLP nonlinearity. We focus on the four dense matrices that dominate the block computation. The fused attention input matrix 𝑊𝑄𝐾𝑉 maps the residual stream into query, key, and value features; in GPT-2 implementations this is often stored as one combined projection and then split into 𝑄, 𝐾, and 𝑉 [21, 32]. The attention output matrix 𝑊𝑂 maps the concatenated head outputs back into the residual stream. The MLP up-projection 𝑊up expands the residual dimension into the hidden MLP dimension, while the MLP down-projection 𝑊down maps the activated hidden representation back into the residual stream. We summarize the component-weight and singular-value associations used in the component-wise interpretation in table 1. When we refer to a block-level aggregate, we combine these component matrices within a Transformer block; when we refer to a component plot, we track one of these matrices across depth. 1 In the later experiments we consider 𝑊 𝑄𝐾𝑉 jointly and 𝑊𝑂 separately. This grouping is a limitation, because the functionality of 𝑊𝑉 is more closely related to 𝑊𝑂 than to 𝑊𝑄 and 𝑊𝐾 .
This measures how many singular directions carry substantial mass [16].
2.3
Pre-trained LLMs Share Similar Spectral Patterns
We analyze eleven Hugging Face checkpoints: English GPT-2 small and medium models [21], Russian GPT-2 small and large models [39], and GPT-2-style Vietnamese, Chinese, Portuguese, Japanese, Turkish, poem-generation, and story-generation models [8, 9, 14, 23, 27, 28, 37, 38]. The selection keeps the architecture family fixed while varying size, language, tokenizer, and training corpus. We analyze block-level aggregates and four subcomponents: the fused 𝑊𝑄𝐾𝑉 attention input projection, the 𝑊𝑂 attention output projection, the 𝑊up MLP up projection, and the 𝑊down MLP down projection. We treat 𝑊𝑄𝐾𝑉 jointly because it determines the query, key, and value feature spaces used to compute and transport attention information. We analyze 𝑊𝑂 separately because it maps attention outputs back into the residual stream, so its spectrum is more directly tied to write-back strength and directionality than to attention-score formation. General Observations. Across pretrained checkpoints, layer-level Frobenius norm and effective-rank entropy follow similar trends after per-model normalization. The shared shape is clearer in zscore plots than metric-value plots, indicating that the common signal is mostly structural rather than a consequence of identical absolute scale. At the component level, effective-rank entropy drops across multiple components in the final few layers. The decrease is most pronounced for the residual-writing or value-side matrices 𝑊𝑂 and 𝑊down , but a similar late-depth decrease also appears for 𝑊𝑄𝐾𝑉 and 𝑊up . This means that the represented spectrum becomes
Complexity-Guided Component-wise Initialization for Language Model Pretraining
more concentrated near the top of the network: fewer singular directions carry a large share of the spectral mass. At the same time, the Frobenius norms of 𝑊𝑂 and 𝑊down rise almost linearly with depth, implying that the total weight mass of these matrices increases while their effective spectra become narrower. The 𝑊𝑄𝐾𝑉 and 𝑊up matrices follow related but weaker depth shapes; common and individual trends are less pronounced there than for 𝑊𝑂 and 𝑊down .
The spectrum-shaping rule separates singular-value shape from total magnitude through the final Frobenius rescaling. Because it uses the normalized singular-value index 𝑖/𝑛, the same rule can be applied across component matrices with different shapes. The resulting initialized profiles reproduce some coarse pretrained tendencies in Frobenius norm and effective-rank entropy. Figure 4 shows these initialization profiles against the pretrained cohort summary.
Interpretation. We interpret the component-level trends through the idea of neural memories [26]. In this view, a two-matrix computation can be read as a key-value memory: one matrix or matrix product determines which features are addressed, while the other determines what information is returned. This interpretation has been applied to Transformer feed-forward layers because their upand down-projections form such a two-matrix computation [5]. We can apply the same key-value interpretation to the attention weight matrices. Under this view, the 𝑊𝑄𝐾 weights function similarly to keys because they determine the attention logits and therefore the attention weights that select or mix token positions [2]. The 𝑊𝑉 𝑂 weights function as the value pathway because they transport the selected content back into the residual stream. This gives a natural interpretation of the weaker Frobenius-norm trend for 𝑊𝑄𝐾𝑉 : changes in attention-logit magnitude are normalized when logits are converted into attention weights explaining the almost constant Frobenius norm, while the late-layer Frobenius-norm growth and effective-entropy drop in 𝑊𝑂 suggest increasingly strong but more spectrally concentrated residual write-back directions. The MLP pattern can then be read through the original up-anddown key-value interpretation: 𝑊up detects or keys features in the hidden MLP space, while 𝑊down writes value-like information back to the residual stream [4, 5]. Because MLP layers are often associated with factual and feature storage [4, 5, 17], the stronger common trends in 𝑊down may reflect later-layer specialization in the readout or residual-writing side of the MLP. Taken together, the strongest shared spectral trends appear in the matrices most directly responsible for writing information into the residual stream, although this should be understood as a qualitative interpretation of the plots rather than a conclusion from a separate quantitative analysis.
3.2
3
Spectral-pattern-based Weight Initialization
We now test whether the spectral patterns identified in pretrained GPT-2-style models can be turned into useful initialization rules for training a new model.
3.1
Proposed Solution
We turn the pretrained depth trends from section 2 into componentwise initialization rules. The proposed initializers are deliberately coarse: they approximate broad scale and spectral concentration profiles rather than fitting every checkpoint-specific curve. This design tests whether pretrained-like scale and singular-value shape are useful initialization signals when the model architecture is fixed. The proposed initializers and comparison baselines are summarized in table 2.
Results
Experiment Setup. All full runs use the same 24-layer GPT-2medium-sized architecture (𝑑 model = 1024, 16 heads, context length 1024) with the GPT-2 tokenizer and SlimPajama-6B [21, 24]. We train for 90k steps, approximately one epoch over 6B tokens, on 8 H20 GPUs with seed 1, per-GPU batch size 2, gradient accumulation 4, learning rate 3 · 10−4 , weight decay 0.1, 10M validation tokens, evaluation every 1000 steps, and checkpoints every 10k steps. We evaluate validation perplexity and held-out perplexity on FineWeb-Edu, OpenWebText, and WikiText-103 [7, 15, 18, 20]. We also evaluate BLiMP syntactic minimal-pair accuracy [31] and multiple-choice accuracy on ARC-Challenge, ARC-Easy, and WinoGrande [1, 22]. We compare the proposed initializers with Standard, High std, No residual scaling, and Pretrained reuse baselines. Empirical Evaluation. The proposed initializers do not show a uniform benefit over the baselines. Magnitude hurts BLiMP accuracy, while Mag.+spectrum has the weakest validation and heldout perplexities. Realistic scale narrows this gap, suggesting that absolute scale is important, but it still does not dominate the Standard baseline. Pretrained reuse is competitive on perplexity, BLiMP, WikiText-103, and ARC-Easy despite the tokenizer/language mismatch, but it is weak on ARC-Challenge. Overall, pretrained spectra are descriptive of trained models, but not sufficient by themselves as an optimization strategy. Table 3 summarizes the evaluation results. Spectral Analysis. The proposed initializers visibly change the structural spectral patterns of the model. The magnitude-based initializers impose clear Frobenius-norm depth profiles at initialization, with decreasing scale for 𝑊𝑄𝐾𝑉 and𝑊up and increasing scale for the residual-writing matrices 𝑊𝑂 and 𝑊down . Mag.+spectrum additionally changes effective-entropy structure by making the initialized spectra more concentrated. After training, these imposed profiles are partly overwritten but not fully erased: the final checkpoints move toward trained-model-like curves while retaining strategydependent differences, especially in component-wise scale and spectral concentration. Figure 5 compares each strategy’s initialization profile with the corresponding trained checkpoint.
4
Discussion
We discuss the main limitations of our spectral analysis and initialization experiments, and outline directions for testing richer forms of pretrained structure reuse.
4.1
Threats to Validity
External Validity. Our analysis is limited to GPT-2-style decoderonly Transformers. This constraint keeps the architecture family
Konstantin Garbers and Nicholas Oh
2
2
1 0
0
1 2
2
3 0.0
0.2
0.4
0.6
0.8
1.0
0.0
(a) Layer Frobenius norm.
0.2
0.4
0.6
0.8
1.0
(b) Layer effective entropy.
Figure 2: Pretrained layerwise Frobenius-norm and effective-entropy trends. 2
4 2
1
0
0
0
0
1
2
2
2
2
2
2 3
0.0
0.2
0.4
0.6
0.8
1.0
0.0
(a) 𝑊𝑄𝐾𝑉 Frobenius norm.
0.2
0.4
0.6
0.8
1.0
0.0
(b) 𝑊𝑄𝐾𝑉 effective entropy.
0.2
0.4
0.6
0.8
1.0
0.0
(c) 𝑊𝑂 Frobenius norm.
0.2
0.4
0.6
0.8
1.0
(d) 𝑊𝑂 effective entropy. 2
2 2
2 0 2
0
1
2
0
0.2
0.4
0.6
0.8
1.0
2 0.0
(e) 𝑊up Frobenius norm.
2
1
4 0.0
0
0.2
0.4
0.6
0.8
1.0
(f) 𝑊up effective entropy.
0.0
0.2
0.4
0.6
0.8
1.0
(g) 𝑊down Frobenius norm.
0.0
0.2
0.4
0.6
0.8
1.0
(h) 𝑊down effective entropy.
Figure 3: Pretrained component-wise Frobenius-norm and effective-entropy trends for attention and MLP projections. Table 2: Initialization methods compared in the training experiments. Type
Baseline
Method
Construction
Standard
Samples linear and embedding weights from a zero-mean Gaussian with standard deviation 0.02 and applies GPT-2 residual-branch scaling [21]. Keeps the standard GPT-2 recipe but increases the initialization standard deviation to 0.08. Uses the standard Gaussian scale but removes GPT-2 residual-branch scaling. Copies non-token-embedding weights from a Chinese GPT-2-medium checkpoint and reinitializes token embeddings for the GPT-2 tokenizer [9].
High std No residual scaling Pretrained reuse Magnitude
Proposed Mag.+spectrum
Realistic scale
Replaces the single global scale with per-layer, per-component target standard deviations: 𝑊𝑄𝐾𝑉 decreases with depth, 𝑊up decreases mildly, and residual-writing projections increase with depth. Uses the same target magnitudes, then tries to match the stronger spectral concentration seen in pretrained models by smoothly decaying the singular values. Concretely, it multiplies singular values by exp(−𝜆(𝑧)𝑖/𝑛), where 𝜆(𝑧) = 5(1 − 𝑧) and 𝑧 is relative depth, before rescaling each √ matrix to ∥𝑊 ∥ 𝐹 = 𝑑 in𝑑 out 𝜎 (𝑧). Applies the magnitude-spectrum construction and then rescales each subcomponent with cohortderived multipliers so the predicted Frobenius RMS matches the trained cohort scale per subcomponent.
fixed and makes component-wise comparisons cleaner, but it also limits how far the conclusions can be transferred to other model families, attention variants, normalization schemes, or larger contemporary LLMs. Within this family, we partially mitigate the limitation by sampling checkpoints with different sizes, languages, tokenizers, and training corpora.
Internal Validity. The training experiments use one main random seed and one primary pretraining dataset. As a result, small differences between initialization strategies may reflect seed-specific optimization noise or dataset-specific effects rather than stable differences in the methods. The strongest negative results are therefore more reliable than fine-grained rankings among runs with similar perplexity or accuracy.
Complexity-Guided Component-wise Initialization for Language Model Pretraining
Effective entropy
Frobenius
z-score
value
z-score
1000
WO
0
500 250
0.0
0.5
1.0
0
100
0.0
0.5
2
500
0 0.5
1.0
1000 0
750
2
500
0.0
250 1.0 0.0
0.5
0.5
1.0
0.0
0.5
1.0
2 0.0
0.5
1.0
Pre-trained
0.5
1.0
Standard
0.0
0.5
1.0
0.0
0.5
1.0
0.5
1.0
400 200
0.0
Magnitude
1.0
100
0
250 0.0
0.5
200
2
500
1
0.0 300
750
0
1.0
100
2 1.0
0.5
200
0
0.5
0 0.0 300
2
1000
1
2 0.0
750
250 1.0 0.0
0.5
1.0
4
0
0.0
Wdown
200
1000
2
Wup
2
750
2
WQKV
value
0.5
Mag.+spectrum
1.0
0 0.0
Realistic scale
Figure 4: Initialization diagnostic grid. Rows correspond to 𝑊𝑂 , 𝑊𝑄𝐾𝑉 , 𝑊up , and 𝑊down . Columns group panels by metric family, effective entropy or Frobenius norm, and by value type, z-score or metric value. In each panel, relative depth is shown on the horizontal axis and the corresponding diagnostic value is shown on the vertical axis. Line styles distinguish the initialization strategies, while the blue dotted curve and shaded band show the pre-trained cohort mean and ±1 standard deviation. Table 3: Evaluation summary grouped by criterion. Arrows indicate whether lower or higher values are better. PPL denotes perplexity, Acc. % denotes BLiMP syntactic minimal-pair accuracy, and MC % denotes multiple-choice accuracy. Bold marks the best result in a column, and red marks clear negative outliers among the full 90k-step runs. Proposed initializers are separated from baselines by a horizontal rule. PPL ↓ Init
Acc. % ↑
MC % ↑
Val
FineWeb
OWT
WikiText
BLiMP
ARC-C
ARC-E
WinoG
Standard High std No residual scaling Pretrained reuse
21.68 22.20 21.95 21.68
29.9 30.7 30.3 29.9
30.3 31.3 31.0 30.7
51.9 51.9 53.2 49.4
80.0 80.9 80.8 80.9
24.2 23.1 23.5 21.8
37.4 38.0 38.0 38.6
50.4 50.3 51.1 50.3
Magnitude Mag.+spectrum Realistic scale
22.06 22.24 21.93
30.5 30.8 30.3
31.2 31.5 31.2
51.3 53.8 52.1
76.2 80.6 80.0
23.3 22.6 22.6
38.1 38.4 37.6
50.6 50.4 50.9
Konstantin Garbers and Nicholas Oh
Effective entropy z-score 2
Standard
Frobenius value
z-score
1000
0 2
250
2
1000
0 2
250
2
1000
0 2
250
2
1000
100
2
50
2
150
0
100
2
50
2
150
0
100
2
50
2
150
0
100
750
0
500 2 0.0
0
750 500
Realistic scale
150
750 500
Mag.+spectrum
2
750 500
Magnitude
value
0.5
1.0
250 0.0
0.5
After training
1.0
2 0.0
0.5
1.0
50 0.0
0.5
1.0
At initialization
Figure 5: Strategy-matched initialization and final-checkpoint diagnostics. Rows correspond to the initialization strategy used for training. Columns follow the same structure as the main initialization grid: effective entropy and Frobenius norm, each shown as z-score and metric value. Dashed colored curves show the closed-form initialization prediction, while solid black curves show the corresponding trained checkpoint. Construct Validity. The proposed initializers approximate broad component-wise spectral trends rather than reproducing every detail of the pretrained spectra. This means they can miss local features such as sharp dips, end-of-depth spikes, or short-range changes that may be functionally important. Our results therefore test whether coarse spectral shape and scale are useful initialization signals, not whether an exact spectral replica of a pretrained model would improve training.
4.2
Future Work
Future work should separate the descriptive role of spectral complexity from its causal effect on downstream performance. In this paper, complexity is partly enforced through initialization, so it remains unclear whether naturally arising changes in effective rank or Frobenius norm predict better generalization when they are not
explicitly imposed. A broader study could measure how these diagnostics correlate with performance across model sizes, datasets, checkpoints, and training stages. Another direction is to study which pretrained patterns are shared across language models and which are specific to a model, language, or corpus. Mechanistic-interpretability work often builds on the linear representation hypothesis, under which high-level features are encoded as directions in model representation spaces; recent work further asks whether such directions can be aligned or transferred across models [12, 13, 19]. Our results suggest an analogous question at the weight level: whether repeated componentwise spectral patterns are merely consequences of architecture and optimization, or whether they encode reusable structure for new training runs. The competitive performance of pretrained-weight reuse, despite a tokenizer and language mismatch, suggests that some form
Complexity-Guided Component-wise Initialization for Language Model Pretraining
of pattern reuse is possible. However, our spectral initializers show that coarse spectrum matching is not enough to capture the benefit. Future initialization methods should test richer reuse signals, such as preserving subspace geometry, component-specific singular vectors, or recurring layer-specific weight patterns, and should evaluate whether any early training advantage persists at longer training horizons.
5
Related Work
This work connects to four lines of prior work. Weight Initialization. Classical initialization methods control activation and gradient scale at the start of training [6, 10]. In Transformer language models, initialization also interacts with residual-depth scaling: GPT-2 scales residual-layer weights at initialization to compensate for accumulation across many residual branches [21]. More recent work studies how initialization scale and spectral control affect training stability and whether learned solutions favor reasoning-like or memorization-like behavior [34– 36]. Our experiments keep the architecture fixed and ask whether pretrained component-wise spectral profiles can provide a useful initialization signal beyond global variance choices in NLP tasks. Spectral Structure in Trained Models. Spectral analysis has been used to characterize implicit regularization and complexity in trained weights [16]. Related analyses of Transformer components show that layers and subcomponents can differ systematically in importance or retained information [25, 29, 33]. Our first-stage analysis follows this diagnostic view, but focuses on whether GPT-2-style checkpoints share recurring component-wise depth profiles across model size, language, tokenizer, and corpus. Mechanistic Interpretability. Mechanistic-interpretability work motivates interpreting Transformer components as structured computations: attention can be decomposed into query-key and outputvalue circuits [2], MLP blocks can behave like key-value memories or vocabulary-space updates [4, 5, 17], and feature directions may be recoverable from singular-vector structure [3, 25]. This motivates our component-wise treatment of 𝑊𝑄𝐾𝑉 , 𝑊𝑂 , 𝑊up , and 𝑊down , although our measurements remain spectral rather than causal or circuit-level. Complexity Control and Generalization. Spectral and low-rank structure is also widely used for model compression and resource allocation in Transformer language models [11, 30]. This line of work treats singular values as a practical handle on model size and retained information, whereas our experiments ask whether pretrained spectral structure can be reused as an initialization signal. Work on cross-model representation alignment suggests that internal feature spaces can sometimes be compared or transferred across models [12, 13, 19]; our setting asks a narrower weight-level version of this question, where only coarse spectral summaries are transferred.
6
Conclusion
We studied whether spectral patterns found in pretrained GPT2-style models can be transferred into the initialization of a new language model. Across eleven pretrained checkpoints, we found
recurring layerwise and component-wise trends despite differences in model size, language, tokenizer, and corpus. The most consistent patterns appear in residual-writing components: 𝑊𝑂 and 𝑊down tend to grow in Frobenius norm with depth while their effective spectra become more concentrated. Our initialization experiments show that these patterns can be imposed structurally. Initializers that replace the single global weight scale with layer- and component-specific magnitudes, and that additionally reshape singular values to mimic the stronger spectral concentration of pretrained models, visibly alter the model’s Frobenius-norm and effective-entropy profiles; some of these differences remain after training. However, these structural changes do not translate into a consistent evaluation gain. The magnitude-only weight initialization intervention hurts BLiMP accuracy, adding singular-value reshaping performs worst on several perplexity measures, and matching pretrained absolute scale more closely narrows but does not remove the gap to standard initialization. By contrast, direct pretrained-weight reuse remains competitive despite a tokenizer and language mismatch, indicating that useful transferable structure may exist but is not captured by coarse spectral summaries alone. The main implication is therefore negative but informative: pretrained spectra describe real regularities in trained Transformer weights, yet copying component-wise scale and singular-value shape is insufficient as a standalone pretraining initialization method. Future work should test richer ways of reusing pretrained structure, such as singular-vector information and component-specific patterns, and should study whether spectral metrics only describe trained models or can directly improve optimization and generalization.
References [1] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv:1803.05457 [cs.AI] https://arxiv.org/abs/1803.05457 [2] Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html [3] Gabriel Franco, Carson Loughridge, and Mark Crovella. 2026. Singular Vectors of Attention Heads Align with Features. arXiv:2602.13524 [cs.LG] doi:10.48550/ arXiv.2602.13524 To be published in ICML 2026. [4] Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022. Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 30–45. doi:10.18653/v1/2022.emnlp-main.3 [5] Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer Feed-Forward Layers Are Key-Value Memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 5484–5495. doi:10.18653/v1/2021.emnlp-main.446 [6] Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 9), Yee Whye Teh and Mike Titterington (Eds.). PMLR, Chia Laguna Resort, Sardinia, Italy, 249–256. https://proceedings.mlr.press/v9/glorot10a.html [7] Aaron Gokaslan and Vanya Cohen. 2019. OpenWebText Corpus. http:// Skylion007.github.io/OpenWebTextCorpus.
Konstantin Garbers and Nicholas Oh
[8] Pierre Guillou. 2020. GPorTuguese-2 (Portuguese GPT-2 small): a Language Model for Portuguese text generation (and more NLP tasks...). https://huggingface.co/pierreguillou/gpt2-small-portuguese. [9] Chengxi Guo. 2023. Mymusise/Gpt2-Medium-Chinese · Hugging Face. https://huggingface.co/mymusise/gpt2-medium-chinese. [10] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision. ICCV, Santiago, Chile, 1026–1034. doi:10.1109/ICCV.2015.123 [11] Ting Hua, Xiao Li, Shangqian Gao, Yen-Chang Hsu, Yilin Shen, and Hongxia Jin. 2023. Dynamic Low-rank Estimation for Transformer-based Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 9275–9287. doi:10.18653/v1/2023.findings-emnlp.621 [12] Youcheng Huang, Chen Huang, Duanyu Feng, Wenqiang Lei, and Jiancheng Lv. 2025. Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 3686–3704. doi:10. 18653/v1/2025.acl-long.185 [13] Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. 2024. Position: The Platonic Representation Hypothesis. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, Vienna, Austria, 20617–20642. https://proceedings.mlr.press/v235/huh24a.html [14] H Toprak Kesgin, M Kaan Yuce, Eren Dogan, M Egemen Uzun, Atahan Uz, H Emre Seyrek, Ahmed Zeer, and M Fatih Amasyali. 2024. Introducing cosmosGPT: Monolingual training for Turkish language models. In 2024 International Conference on INnovations in Intelligent SysTems and Applications (INISTA). IEEE, INISTA, Craiova, Romania, 1–6. [15] Anton Lozhkov, Loubna Ben Allal, Leandro von Werra, and Thomas Wolf. 2024. FineWeb-Edu: the Finest Collection of Educational Content. doi:10.57967/hf/2497 [16] Charles H Martin and Michael W Mahoney. 2021. Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning. Journal of Machine Learning Research 22, 165 (2021), 1–73. [17] Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and Editing Factual Associations in GPT. Advances in Neural Information Processing Systems 35 (Dec. 2022), 17359–17372. [18] Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer Sentinel Mixture Models. arXiv:1609.07843 [cs.CL] [19] Narmeen Fatimah Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Michael Lan, Abir Harrasse, and Amir Abdullah. 2025. Activation space interventions can be transferred between large language models. In Proceedings of the 42nd International Conference on Machine Learning (Vancouver, Canada) (ICML’25). JMLR.org, Vancouver, Canada, Article 1890, 75 pages. [20] Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Associates, Inc., Vancouver, Canada, 30811–30849. doi:10.52202/079017-0970 [21] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9. [22] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. WinoGrande: An Adversarial Winograd Schema Challenge at Scale. arXiv:1907.10641 [cs.CL] https://arxiv.org/abs/1907.10641 [23] Kei Sawada, Tianyu Zhao, Makoto Shing, Kentaro Mitsui, Akio Kaga, Yukiya Hono, Toshiaki Wakatsuki, and Koh Mitsuda. 2024. Release of Pre-Trained Models for the Japanese Language. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LRECCOLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 13898–13905. https://aclanthology.org/2024.lrec-main.1213/ [24] Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey. 2023. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. https://www.cerebras.net/blog/slimpajama-a-627btoken-cleaned-and-deduplicated-version-of-redpajama. https://huggingface. co/datasets/cerebras/SlimPajama-627B [25] Max Staats, Matthias Thamm, and Bernd Rosenow. 2026. Small singular values matter: A random matrix analysis of transformer models. Advances in Neural Information Processing Systems 38 (2026), 153545–153573. [26] Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. 2015. Endto-end memory networks. In Proceedings of the 29th International Conference on Neural Information Processing Systems - Volume 2 (Montreal, Canada) (NIPS’15). MIT Press, Cambridge, MA, USA, 2440–2448.
[27] Pranav Vadrevu. 2021. Pranavpsv/Gpt2-Genre-Story-Generator · Hugging Face. https://huggingface.co/pranavpsv/gpt2-genre-story-generator. [28] Nha Nguyen Van. 2025. NlpHUST/Gpt2-Vietnamese · Hugging Face. https://huggingface.co/NlpHUST/gpt2-vietnamese. [29] Wenxuan Wang and Zhaopeng Tu. 2020. Rethinking the Value of Transformer Components. In Proceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on Computational Linguistics, Barcelona, Spain (Online), 6019–6029. doi:10.18653/v1/2020.coling-main.529 [30] Xin Wang, Samiul Alam, Zhongwei Wan, Hui Shen, and Mi Zhang. 2025. SVDLLM V2: Optimizing Singular Value Truncation for Large Language Model Compression. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Luis Chiruzzo, Alan Ritter, and Lu Wang (Eds.). Association for Computational Linguistics, Albuquerque, New Mexico, 4287– 4296. doi:10.18653/v1/2025.naacl-long.217 [31] Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, ShengFu Wang, and Samuel R. Bowman. 2020. BLiMP: The Benchmark of Linguistic Minimal Pairs for English. Transactions of the Association for Computational Linguistics 8 (2020), 377–392. arXiv:https://doi.org/10.1162/tacl_a_00321 doi:10. 1162/tacl_a_00321 [32] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Qun Liu and David Schlangen (Eds.). Association for Computational Linguistics, Online, 38–45. doi:10.18653/v1/2020.emnlp-demos.6 [33] Lu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh, Yaqing Wang, Yiling Jia, Gen Li, Ajay Jaiswal, Mykola Pechenizkiy, Yi Liang, Michael Bendersky, Zhangyang Wang, and Shiwei Liu. 2025. Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity. arXiv:2310.05175 [cs.LG] https://arxiv.org/abs/2310.05175 [34] Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang, Jiatao Gu, and Josh Susskind. 2023. Stabilizing transformer training by preventing attention entropy collapse. In Proceedings of the 40th International Conference on Machine Learning (ICML’23). JMLR.org, Honolulu, Hawaii, USA, Article 1709, 34 pages. [35] Zhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang, and Zhi-Qin John Xu. 2026. Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence 48, 4 (2026), 4336–4349. doi:10.1109/TPAMI.2025.3646483 [36] Zhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang, and Zhi-Qin John Xu. 2024. Initialization is critical to whether transformers fit composite functions by reasoning or memorizing. In Proceedings of the 38th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS ’24). Curran Associates Inc., Red Hook, NY, USA, Article 451, 34 pages. [37] Tianyu Zhao and Kei Sawada. 2021. rinna/japanese-gpt2-medium. https:// huggingface.co/rinna/japanese-gpt2-medium [38] Zhe Zhao, Yudong Li, Cheng Hou, Jing Zhao, Rong Tian, Weijie Liu, Yiren Chen, Ningyuan Sun, Haoyan Liu, Weiquan Mao, et al. 2023. Tencentpretrain: A scalable and flexible toolkit for pre-training models of different modalities. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). ACL, Toronto, Canada, 217–225. [39] Dmitry Zmitrovich, Aleksandr Abramov, Andrey Kalmykov, Vitaly Kadulin, Maria Tikhonova, Ekaterina Taktasheva, Danil Astafurov, Mark Baushenko, Artem Snegirev, Tatiana Shavrina, Sergei S. Markov, Vladislav Mikhailov, and Alena Fenogenova. 2024. A Family of Pretrained Transformer Language Models for Russian. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 507–524. https://aclanthology.org/2024.lrec-main.45/