Conceptio › Archive › arXiv CS
arXiv CSopen access

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Preprint.

H OW D OES mHC U SE I TS R ESIDUAL S TREAMS ? S E LECTIVE ROUTING AND N EAR -I DENTITY M IXING Pengxiang Zhao1

Xing Li1

Xianzhi Yu1

Wei Guo1

Zhenhua Dong1

1 Huawei Technologies Co., Ltd.

{zhao.pengxiang, li.xing2}@huawei.com

arXiv:2609.05309v1 [cs.LG] 4 Sep 2026

A BSTRACT Hyper-Connections and their manifold-constrained variant (mHC) widen a residual pathway from one stream to 𝑛, yet how trained models use this capacity remains unclear: how broadly blocks read and write, how strongly the residual pathway mixes streams, and whether the streams carry distinct representations. We examine these properties in the four-stream residual pathway of DeepSeek-V4-Flash using effective stream counts, cross-stream residual weights, and inter-stream cosine similarity. Read/write routing is concentrated but varies across depth: a typical attention or FFN site effectively uses about two streams, while the dominant stream changes across layers and the representations remain directionally distinct. Residual mixing is modest and occurs primarily in early layers; in layers 22–42, the pathway mostly carries each stream forward separately. Targeted interventions establish the functional significance of these patterns. Replacing the late mixers by identity increases C4 perplexity by only 1.9% and preserves the six-task average score, whereas replacing the early mixers increases perplexity by 41%. Fixing each early mixer to its C4 diagnostic mean increases perplexity by only 0.2% and reduces the average score by 0.25 percentage points, showing that its site-specific structure matters more than its token-wise variation on the evaluated metrics. Likewise, retaining the three largest routing weights per token at every site increases perplexity by at most 2.7% and changes the average score by at most 0.4 points. Thus, the studied model realizes only part of the flexibility afforded by four-stream mHC: individual blocks rarely require all four streams, and late residual mixing provides little measured benefit.

Layer read

0

10

F

write

L1 Attn

20

30

42

Late (22--42)

Early (0--21)

F

F

L13 FFN

L28 Attn

F

L40 FFN

Stream 0 Stream 1 Stream 2 Stream 3 Set late H res to I4

Top-3 routing

read: ΔPPL +2.4%, Δacc. −0.38 pp write: ΔPPL +2.7%, Δacc. −0.05 pp

ΔPPL +1.9%, Δacc. +0.04 pp

Figure 1: Realized multi-stream computation in DeepSeek-V4-Flash. At four sites spanning depth and both sublayer types, color identifies the source stream, arrows show the direction of information flow, and line width encodes the measured token-averaged routing or residual weight. Routing concentrates on changing stream subsets, whereas residual exchange largely disappears after mid-depth; the callouts give the exact intervention effects on C4 perplexity and the six-task average score.

1

Preprint.

1

I NTRODUCTION

Hyper-Connections (HC) (Zhu et al., 2024) generalize the standard single-stream residual pathway (He et al., 2016) to 𝑛 coupled streams. At each sublayer, a read map aggregates the streams into a single branch input, a write map distributes the branch output back across the streams, and a residual mixer linearly recombines the incoming stream states along the skip pathway. Manifoldconstrained Hyper-Connections (mHC) (Xie et al., 2025) are designed to improve training stability by constraining each residual mixer to be doubly stochastic, thereby limiting signal amplification as these mixers compose across depth. This design widens the state propagated across depth while retaining a single computation branch at each sublayer. However, mHC makes broad routing and cross-stream mixing possible, while training determines how and where the resulting model uses this flexibility. Because the read, write, and residual maps are token-conditioned and learned at each layer, their realized behavior can vary across both tokens and depth. Accordingly, the read and write maps may distribute branch computation broadly or concentrate it on a subset of streams; the residual pathway may mix the stream states or propagate them separately; and the streams may encode distinct or redundant representations. Recent work on 120M- and 360M-parameter HC models built with nanoGPT (Karpathy, 2022) finds a dominant-stream regime: read/write signals and interpretable features concentrate in one stream, while residual mixing remains close to identity (Alimaskina et al., 2026). Because model scale and training regime may shape how the learned maps allocate computation across streams, it remains unclear whether this regime is characteristic of trained mHC models or only one of several possible organizations of multi-stream computation. Moreover, these structural patterns do not by themselves establish post-training functional redundancy. Determining this requires testing whether weak routing weights can be pruned and whether near-identity residual mixers can be replaced by exact identity maps without degrading model quality. We address both gaps with a true-forward analysis of DeepSeek-V4-Flash, a 284B-parameter mHC model pretrained on more than 32T tokens (DeepSeek-AI, 2026). Our framework measures effective read/write width, cross-stream residual weights, and inter-stream representation similarity across tokens, sublayers, and depth. These depth-resolved diagnostics distinguish a single globally dominant stream from shifting stream use across layers and determine whether the streams retain distinct representations. We pair them with targeted inference-time interventions to connect the observed routing and mixing patterns to model quality. Applying this framework reveals concentrated but depth-varying routing: a typical attention or FFN site effectively uses about two of the four streams, while the dominant stream changes across layers and the stream representations remain directionally distinct. Routing ablations show that this concentration creates local redundancy but does not reduce the pathway to two streams. Retaining the three largest realized read or write weights for every token at every site increases perplexity on C4 (Raffel et al., 2020) by at most 2.7% and changes the average downstream score by at most 0.4 percentage points; retaining only two weights per token instead increases perplexity by 12–14% and reduces the average downstream score by 2.9–6.6 percentage points. Residual mixing also varies with depth. Cross-stream residual weights are modest overall and most pronounced in layers 0–21 of the 43-layer model, whereas the residual mixers in layers 22–42 remain close to the identity. Residual-mixer ablations confirm that this depth profile has functional significance. Replacing the mixers in layers 22–42 by exact identity maps increases C4 perplexity by only 1.9% and preserves the average score across six downstream benchmarks, whereas replacing those in layers 0–21 increases perplexity by 41%. Yet fixing each early mixer to its C4 diagnostic mean raises perplexity by only 0.2% and reduces the six-task average score by 0.25 points, indicating that its learned site-specific structure accounts for most of its measured benefit, while token-wise variation contributes little on the evaluated metrics. Figure 1 provides a data-driven overview of this analysis, tracing realized routing and residual exchange across representative depths and linking both to their measured intervention effects. In summary, our contributions are threefold: • True-forward analysis of multi-stream computation. We introduce a framework that resolves read/write routing, residual mixing, and stream representations across tokens, sub2

Preprint.

WRITE

READ pre 𝑯𝑙

incoming streams

1× 𝑛

pre

𝑯𝑙 𝑿𝑙

post ⊤ 𝑛 ×𝐶 )

(𝑯𝑙

F (·, 𝑾𝑙 )

𝑛 ×1

outgoing streams

(1) 𝒙𝑙

(1)

+

. . . (𝑛) 𝒙𝑙

𝑛 ×𝐶

𝑯𝑙res

𝒙𝑙+1 . . . (𝑛) 𝒙𝑙+1

𝑛×𝑛

RESIDUAL MIXER pre

Figure 2: Data flow through HC residual sublayer 𝑙. The read map 𝑯𝑙 forms the computationpost branch input, the write map 𝑯𝑙 distributes the branch output across streams, and the residual res mixer 𝑯𝑙 carries forward and linearly recombines the incoming stream states. layers, and depth, and connects these measurements to model quality through inferencetime interventions. • A depth-resolved characterization of DeepSeek-V4-Flash. We show that its four streams exhibit locally concentrated but depth-varying routing, retain distinct representations, and exchange information primarily in the first half of the network. • Functional localization of routing and mixing capacity. Controlled interventions identify where additional routing weights and learned residual mixing affect model quality, revealing local redundancy in read/write maps and limited functional dependence on latelayer residual mixing.

2

BACKGROUND AND R ELATED W ORK

2.1

R ESIDUAL S TREAMS AND H YPER -C ONNECTIONS

Consider a residual sublayer with hidden dimension 𝐶. A standard residual architecture propagates a single state, whereas Hyper-Connections (HC) maintain 𝑛 coupled states 𝑿𝑙 ∈ R𝑛×𝐶 at sublayer 𝑙 (Zhu et al., 2024). Here and below, 𝑙 indexes residual sublayers, including the attention and FFN sublayers within each Transformer layer; we suppress the token index for clarity. Each sublayer pre post applies three maps to this state: a read map 𝑯𝑙 ∈ R1×𝑛 , a write map 𝑯𝑙 ∈ R1×𝑛 , and a residual res 𝑛×𝑛 mixer 𝑯𝑙 ∈ R . In HC, each map combines learned static parameters with a token-dependent component. We refer to the resulting value for a given token and sublayer as the realized map. Figure 2 summarizes the resulting data flow. Let F (·, 𝑾𝑙 ) denote the computation performed by the sublayer. Its multi-stream update is post ⊤

𝑿𝑙+1 = 𝑯𝑙res 𝑿𝑙 + (𝑯𝑙

pre

) F (𝑯𝑙 𝑿𝑙 , 𝑾𝑙 ),

(1)

pre post pre where 𝑯𝑙 𝑿𝑙 forms the input to the computation branch, (𝑯𝑙 ) ⊤ F (𝑯𝑙 𝑿𝑙 , 𝑾𝑙 ) distributes its output across the streams, and 𝑯𝑙res 𝑿𝑙 carries forward and linearly recombines the incoming stream

states. 2.2

M ANIFOLD -C ONSTRAINED H YPER -C ONNECTIONS

Across successive residual sublayers, the residual mixers compose multiplicatively along the skip pathway. For sublayers 𝑙, . . . , 𝐿 − 1, the resulting cumulative mixer is res res := 𝑯 res Π 𝐿:𝑙 𝐿−1 · · · 𝑯𝑙 .

(2)

Without constraints on the individual mixers, the norm of this product can grow rapidly with depth, amplifying signals propagated along the residual pathway. To control this cumulative transformation, mHC constrains every realized residual mixer to the Birkhoff polytope (Xie et al., 2025):  ⊤ B𝑛 := 𝑯 ∈ R𝑛×𝑛 𝑯1𝑛 = 1𝑛 , 1⊤ (3) 𝑛 𝑯 = 1𝑛 , 𝑯 ≥ 0 , 3

Preprint.

where 1𝑛 is the all-ones vector. In practice, mHC implements this projection using Sinkhorn–Knopp normalization (Sinkhorn & Knopp, 1967; Xie et al., 2025). Under the exact constraint, the Birkhoff polytope is closed under multiplication, and every 𝑯 ∈ B𝑛 has spectral norm one. Consequently, res remains non-expansive, preventing signal amplification along the skip the cumulative mixer Π 𝐿:𝑙 pathway. Although 𝑰𝑛 is a valid doubly stochastic mixer, mHC learns 𝑯𝑙res rather than fixing it to 𝑰𝑛 to enable cross-stream information exchange (Xie et al., 2025). 2.3

A NALYSIS OF L EARNED S TREAM U TILIZATION

The original HC and mHC studies establish the optimization and performance properties of multistream residual pathways (Zhu et al., 2024; Xie et al., 2025). Recent analysis of 120M- and 360Mparameter HC models trained with nanoGPT further identifies a dominant-stream regime in which read/write signals and interpretable features concentrate in one stream while the residual mixers remain close to identity (Karpathy, 2022; Alimaskina et al., 2026). Breaking initialization symmetry can mitigate this concentration (Alimaskina et al., 2026). These findings establish stream collapse as one possible organization of learned HC computation. Our study examines how routing, residual mixing, and stream representations are organized in a separately trained large-scale mHC model and tests the functional significance of these patterns through inference-time interventions.

3

M EASURING R EALIZED M ULTI -S TREAM C OMPUTATION pre

post

For each token at residual sublayer 𝑙, the realized maps 𝑯𝑙 , 𝑯𝑙 , and 𝑯𝑙res determine how branch computation is routed and how residual states are mixed across streams. We characterize this computation along three axes: the breadth of read/write routing, the amount of cross-stream residual mixing, and the similarity between stream representations. 3.1

R EAD /W RITE ROUTING B READTH

We characterize read/write routing at each sublayer from two complementary perspectives: whether tokens share the same dominant stream and how broadly the routing weights are distributed across 𝑞 streams. For diagnostic sequence 𝑚, let ℎ𝑙,𝑚,𝑡 , 𝑗 denote the realized routing weight assigned to stream 𝑗 for its 𝑡-th token at sublayer 𝑙, where 𝑞 ∈ {pre, post} and 𝑗 ∈ {1, . . . , 𝑛}. The realized read and write weights are nonnegative under the mHC parameterization, so we normalize them to obtain the token-level load distribution 𝑞 ℎ𝑙,𝑚,𝑡 ,𝑗 𝑞 := 𝑝 𝑙,𝑚,𝑡 , 𝑗 . (4) Í𝑛 𝑞 ℎ 𝑘=1 𝑙,𝑚,𝑡 ,𝑘 We summarize the resulting token-level distributions with two sublayer-level statistics. Winner consistency. For each token, we identify the stream receiving the largest routing weight. For an input sequence 𝑚 containing 𝑇𝑚 tokens, its winner consistency is the largest fraction of tokens sharing the same winner:  𝑇𝑚  1 ∑︁ 𝑞 𝑞 := max (5) 𝑐 𝑙,𝑚 I 𝑗 = arg max 𝑝 𝑙,𝑚,𝑡 ,𝑘 , 𝑘 𝑗 ∈ {1,...,𝑛} 𝑇𝑚 𝑡=1 𝑞 𝑞 1 Í𝑀 where 𝑐 𝑙,𝑚 ∈ [1/𝑛, 1]. We report 𝑐 𝑙𝑞 = 𝑀 𝑚=1 𝑐 𝑙,𝑚 over the 𝑀 diagnostic sequences. A value of 1 means that all tokens within every sequence share the same winner, while lower values indicate that the dominant stream varies across tokens. Load-effective stream count. quence:

We first aggregate the normalized routing weights within each se𝑇

𝑞 := 𝑝¯𝑙,𝑚, 𝑗

𝑚 1 ∑︁ 𝑝𝑞 . 𝑇𝑚 𝑡=1 𝑙,𝑚,𝑡 , 𝑗

The effective number of streams carrying this load is 1 𝑞 𝑛eff ( 𝒑¯ 𝑙,𝑚 ) := Í𝑛 ∈ [1, 𝑛]. 𝑞 2 𝑗=1 ( 𝑝¯𝑙,𝑚, 𝑗 ) 4

(6)

(7)

Preprint.

𝑘 = 3, 𝛼 = .91−1 0.41

0.29

0.21

0.09

𝑘 = 2, 𝛼 = .70 0.41

0.29

0.21

0.45

0.32

0.23

0

0.59

0.41

0

0

−1

0.09

Figure 3: Scale-preserving top-𝑘 routing for 𝑘 ∈ {2, 3}. Retained weights are rescaled by 𝛼 to preserve their original sum. The effective stream count equals 1 when all routing load falls on one stream and 𝑛 when the load is uniform. As with winner consistency, we report this quantity averaged over the 𝑀 diagnostic sequences. 3.2

C ROSS -S TREAM R ESIDUAL M IXING

We quantify how strongly each residual sublayer mixes its incoming streams by measuring the deres ∈ R𝑛×𝑛 denote the realized residual viation of its realized residual mixer from the identity. Let 𝑯𝑙,𝑡 mixer for token 𝑡 at sublayer 𝑙. We define its identity deviation as 𝑛

res 𝛿(𝑯𝑙,𝑡 ) :=

𝑛

1 ∑︁ ∑︁ res (𝑯𝑙,𝑡 )𝑖 𝑗 − I[𝑖 = 𝑗] . 𝑛2 𝑖=1 𝑗=1

(8)

The metric is zero when each stream is carried forward independently through the identity mixer. Under the exact doubly stochastic constraint, it is proportional to the total off-diagonal weight and therefore increases with cross-stream residual mixing. We average identity deviation across tokens within each sublayer and examine its profile across depth. 3.3

I NTER -S TREAM R EPRESENTATION S IMILARITY

We assess whether the residual streams carry distinct representations by measuring their pairwise (𝑠) alignment. Let 𝒙𝑙,𝑡 ∈ R𝐶 denote the hidden state carried by stream 𝑠 for token 𝑡 at sublayer 𝑙. We define the mean pairwise cosine similarity as ( 𝑗) (𝑖) ∑︁ ⟨𝒙𝑙,𝑡 , 𝒙𝑙,𝑡 ⟩ 2 𝑠𝑙,𝑡 := . 𝑛(𝑛 − 1) 1≤𝑖< 𝑗 ≤𝑛 ∥𝒙 (𝑖) ∥ 2 ∥𝒙 ( 𝑗 ) ∥ 2 𝑙,𝑡

(9)

𝑙,𝑡

Higher values indicate more closely aligned stream representations, whereas lower values indicate greater directional separation. We average this similarity across tokens within each sublayer and examine how it evolves across depth. 3.4

I NFERENCE -T IME I NTERVENTIONS

The diagnostics above describe how the trained model uses its residual streams, but establishing which components are functionally necessary requires a counterfactual test: would model quality change if selected routing paths or cross-stream residual interactions were suppressed? We perform this test through two targeted inference-time interventions without retraining: routing sparsification and residual-mixer replacement. For routing sparsification, we suppress the sequence index 𝑚 and let T𝑙,𝑡𝑞 (𝑘) index the 𝑘 largest realized weights of the read or write map for token 𝑡 at sublayer 𝑙. We define the scale-preserving top-𝑘 map as Í𝑛 𝑞 𝑟=1 ℎ𝑙,𝑡 ,𝑟 𝑞 , (10) 𝛼𝑙,𝑡 (𝑘) := Í 𝑞 𝑞 𝑟 ∈ T𝑙,𝑡 (𝑘 ) ℎ𝑙,𝑡 ,𝑟 h i 𝑞 𝑞 𝑞 𝑞 e := 𝛼𝑙,𝑡 ℎ𝑙,𝑡 (𝑘) ℎ𝑙,𝑡 (11) ,𝑗 , 𝑗 I 𝑗 ∈ T𝑙,𝑡 (𝑘) .

5

Preprint.

This intervention changes the routing support while preserving the total routing weight. Figure 3 illustrates the operation for 𝑘 ∈ {2, 3}. We apply it to the read and write maps separately. res at every token within a selected depth For residual-mixer replacement, we substitute 𝑰𝑛 for 𝑯𝑙,𝑡 range. This removes learned cross-stream exchange within that range while leaving the per-stream skip paths unchanged. We evaluate both interventions using perplexity and downstream-task scores.

4

E XPERIMENTS

We apply the framework from Section 3 to DeepSeek-V4-Flash. We first describe the model, data, and true-forward collection protocol. We then examine read/write routing, stream representations, and residual mixing across depth, followed by inference-time interventions on the learned routing and residual maps. 4.1

E XPERIMENTAL S ETUP

Model. We study the publicly released 0731 checkpoint of DeepSeek-V4-Flash (DeepSeek-AI, 2026), a 43-layer Mixture-of-Experts (MoE) language model with 284B total and 13B activated parameters, pretrained on more than 32T tokens. Its 4096-dimensional backbone combines hybrid attention and MoE feed-forward networks with four-stream mHC residual pathways. We instrument the attention and MoE-FFN residual sublayers in every layer, yielding 86 observation sites. Data and evaluation. We use C4 (Raffel et al., 2020) following the sampling and evaluation protocols of GPTQ (Frantar et al., 2023). For the true-forward diagnostics, we randomly sample 512 sequences of 1,024 tokens from the first English C4 training shard. For perplexity, we follow the GPTQ evaluation protocol (Frantar et al., 2023), evaluating 256 contiguous sequences of 2,048 tokens constructed from the first English C4 validation shard. We additionally evaluate six downstream benchmarks in the zero-shot setting: ARC-Easy and ARC-Challenge (Clark et al., 2018), PIQA (Bisk et al., 2020), HellaSwag (Zellers et al., 2019), MMLU (Hendrycks et al., 2021), and GSM8K (Cobbe et al., 2021). We report normalized accuracy for ARC-Easy, ARC-Challenge, PIQA, and HellaSwag, accuracy for MMLU, and exact-match accuracy for GSM8K, together with the unweighted average of the six task scores. Implementation. We build our forward-pass instrumentation and inference-time interventions on PyTorch (Paszke et al., 2019) and Hugging Face Transformers (Wolf et al., 2020). Downstream evaluation uses the LM Evaluation Harness (Gao et al., 2023). 4.2

O RGANIZATION OF R EAD /W RITE ROUTING

We first characterize how the trained model routes branch computation across its four residual streams. Figure 4 shows that routing is both narrow and largely stable across tokens at a given sublayer. Winner consistency averages 0.871 for the read maps and 0.905 for the write maps, with medians of 0.931 and 0.988 across sublayers; thus, most maps assign a large majority of tokens within an input sequence to the same dominant stream. The corresponding load-effective stream counts average 1.998 and 1.775, with medians of 1.876 and 1.542. A typical sublayer therefore concentrates most of its routing load on roughly two streams, with stronger concentration in the write maps. This local concentration does not reduce to a single stream dominating throughout the network. Figure 5 instead reveals depth-structured assignments: for a fixed map and sublayer type, 62.5% of adjacent layer pairs retain the same winner, producing contiguous intervals whose dominant stream changes at several depth boundaries. The reorganization is especially clear in the read maps. Streams 0 and 1 dominate 32 of the 44 sites in layers 0–21, whereas streams 2 and 3 dominate 36 of the 42 sites in layers 22–42. Every stream therefore becomes dominant in read or write routing somewhere in the network, but over different depth ranges. Routing breadth also changes with depth, primarily through the write maps. Their mean effective stream count increases from 1.467 in layers 0–21 to 2.098 in layers 22–42, while winner consistency 6

Preprint.

1.0

Load effective streams

Winner consistency

0.9 0.8 0.7 0.6 0.5

Hpre Hpost

0.4 0

10

20

30

4.0

Hpre

3.5

Hpost

3.0 2.5 2.0 1.5 1.0 0

40

10

20

30

40

Layer

Layer

(a) Winner consistency.

(b) Load-effective stream count.

Figure 4: Depth profiles of routing concentration. Each point averages the attention and FFN sites within a layer. (a) Winner consistency measures the fraction of tokens sharing the most frequent dominant stream; its read/write averages are 0.871 and 0.905. (b) The load-effective stream count averages 1.998 and 1.775, respectively, indicating that routing typically spans only about two of the four streams. The dashed line in (b) marks two effective streams.

0.2

10

0.0

0

20

0.4 0.2

10

0.0

1

2

3

0.6 20

0.4 0.2

10

0 0

0.8

0.0

0 0

1

2

3

1.0 30

0.8 0.6

20

0.4 0.2

10

0.0

Mean stream load

0.4

0.6

30

Layer

20

0.8

Layer

0.6

30

40 1.0

Mean stream load

0.8

Layer

Layer

30

40 1.0

Mean stream load

40 1.0

Mean stream load

40

0 0

1

2

3

0

1

2

3

Stream

Stream

Stream

Stream

(a) 𝑯 pre at attention sites.

(b) 𝑯 pre at FFN sites.

(c) 𝑯 post at attention sites.

(d) 𝑯 post at FFN sites.

Figure 5: Mean normalized routing load by layer and stream, shown separately for the read and write maps at attention and FFN sites. High-load regions form contiguous depth intervals, but their stream assignments change across layers, sublayer types, and read/write maps. decreases from 0.940 to 0.870; read width remains nearly unchanged at 1.971 and 2.025, respectively. The read and write schedules are coupled but not identical: their winners differ at 38.4% of the 86 sites, so a sublayer need not write primarily to the stream from which it reads. Taken together, routing is locally narrow and largely token-stable, yet its allocation is reorganized across depth rather than collapsing onto one globally dominant stream. Appendix A.1 reports the routing statistics by depth range and sublayer type, while Appendix A.2 identifies the lowest-load stream at each site. 4.3

I NTER -S TREAM R EPRESENTATION G EOMETRY

Having established that routing is locally concentrated, we ask whether this concentration is accompanied by redundancy in the stream representations. Figure 6 tracks their pairwise cosine similarity across depth. The streams are identical at the first attention site, where their mean pairwise cosine is 1.0, but separate rapidly as the first layers transform and redistribute their states. From layer 1 onward, mean similarity ranges from 0.235 to 0.564 across attention and FFN sites, with an average of 0.404. Their directional separation is not monotonic: the streams partially realign around layers 10–13, separate more strongly around layers 27–29, and become progressively more aligned toward the output, without returning to their initial agreement. The pair-resolved view further shows that this geometry is heterogeneous rather than symmetric. Beyond layer 0, pair-specific averages range from 0.288 for streams 0 and 3 to 0.560 for streams 1 and 2, and no pair remains uniformly aligned across depth. Thus, local routing concentration does not 7

Preprint.

Attn FFN

1.0 s0 --s1

0.8

s0 --s2 s0 --s3

0.6

0.5 s1 --s2

0.4

Cosine

Mean pairwise cosine

1.0

s1 --s3 s2 --s3

0.2 0

10

20

30

40

0

10

Layer

20

30

40

0.0

Layer

(a) Mean over the six stream pairs.

(b) Pair-resolved similarity.

Figure 6: Inter-stream representation similarity across depth. (a) Pairwise cosine similarity, averaged over tokens and the six stream pairs at each attention and FFN site. (b) Token-averaged cosine similarity for each stream pair, with attention and FFN sites shown in forward-pass order within each layer. The streams begin from identical states but rapidly diverge, and no stream pair remains uniformly aligned throughout the network.

0

0.15

0.92

0.02

0.05

0.02

1.0

0

0.98

0.01

0.01

0.00

1

0.00

0.98

0.00

0.00

2

0.01

0.00

0.97

0.00

3

0.01

0.01

0.01

1.00

0

1

2

3

0.05

0.02

0.89

0.08

0.01

2

0.04

0.09

0.86

0.01

0.6

(H res )ij

0.10

1

0.4 0.2

3

0.02

0.01

0.01

0.96

0

1

2

3

0.00 0

10

20

30

40

Target stream

Layer

(a) Identity deviation across depth.

0.0

(b) Layers 0–21; 𝛿 = 0.046.

1.0 0.8

Source stream

0.8

0.6

(H res )ij

Attn FFN

Source stream

Identity deviation

0.20

0.4 0.2 0.0

Target stream

(c) Layers 22–42; 𝛿 = 0.009.

Figure 7: Depth structure of realized residual mixing. (a) Identity deviation at each attention and FFN site, averaged over diagnostic tokens. (b–c) Residual mixers averaged over tokens and sites in the indicated depth ranges. Off-diagonal transfer is localized primarily to early layers, while the later mean mixer is nearly diagonal and performs little direct cross-stream exchange.

coincide with representation collapse: even weakly routed streams can carry directionally distinct states. However, representational distinctness does not establish functional importance; Section 4.5 tests whether these weak routing paths materially affect model quality. 4.4

D EPTH S TRUCTURE OF R ESIDUAL M IXING

We next examine how the strength of learned residual mixing varies across depth. Figure 7a reveals pronounced variation across the network: identity deviation averages 0.028 across all 86 sites, but the largest deviations occur at only a small number of early sites. After a final attention-side spike at layer 22, the mixers settle close to identity: 32 of the 40 sites in layers 23–42 have identity deviation below 0.01. Figures 7b and 7c show how this depthwise reduction appears in the residual mixers themselves. Over layers 0–21, the mean mixer retains visible off-diagonal weights, with an identity deviation of 0.046. Over layers 22–42, these weights largely vanish and the deviation falls to 0.009, leaving a near-diagonal map that carries each stream forward with little direct residual exchange. The complete layerwise matrices for attention and FFN sites are provided in Appendix B (Figures 9 and 10). The late mixers therefore realize little of their available cross-stream flexibility despite being learned and token-dependent by design. Their near-static identity structure motivates replacing them with 𝑰4 8

Preprint.

at inference, thereby avoiding the corresponding Sinkhorn–Knopp iterations; Section 4.5 evaluates the resulting effect on model quality. 4.5

F UNCTIONAL E FFECTS OF ROUTING AND M IXING I NTERVENTIONS

The preceding diagnostics characterize realized stream use, but they do not determine which routing and mixing effects are necessary for model quality. We therefore apply the inference-time interventions from Section 3.4, without retraining, and measure their effects on C4 perplexity and downstream-task scores in Tables 1 and 2. Table 1: C4 perplexity under the inference-time interventions. “Token mean” replaces each selected mixer by its average over the C4 diagnostic sequences. Type

Action

Routing

Mixer

PPL ↓

Δ PPL

Relative Δ

11.1587

—

—

+0.2633 +0.3000 +1.3465 +1.5619

+2.4% +2.7% +12.1% +14.0%

+0.2141 +4.6248 +4.7255 +0.0245

+1.9% +41.4% +42.3% +0.2%

Scope

Baseline top-3

𝑯 pre 𝑯 post

top-2

𝑯 pre 𝑯 post

11.4220 11.4587 12.5052 12.7206

22–42 0–21 All layers 0–21

11.3728 15.7835 15.8842 11.1832

𝑯 res → 𝑰4 Token mean

Table 2: Zero-shot downstream performance under the inference-time interventions. All values are percentages; Δ Avg. is measured in percentage points. Type

Action

Scope

Baseline Routing

top-3 top-2

Mixer

𝑯 pre 𝑯 post 𝑯 pre 𝑯 post

22–42 All layers Token mean 0–21 𝑯 res → 𝑰4

Δ Avg.

ARC-E ARC-C PIQA HellaSwag MMLU GSM8K

Avg.

87.92

66.38

84.22

85.63

85.44

95.75

84.22

—

87.96 87.84 85.02 87.42

66.30 66.98 62.54 66.81

83.90 84.22 82.81 82.97

85.39 85.85 83.96 84.99

85.07 85.37 81.66 85.41

94.39 94.77 91.81 58.15

83.84 84.17 81.30 77.63

-0.38 -0.05 -2.92 -6.59

88.05 85.56 87.71

67.32 62.71 67.49

84.28 83.30 83.51

85.81 82.98 85.61

85.41 79.90 85.14

94.69 91.21 94.39

84.26 80.94 83.98

+0.04 -3.28 -0.25

Routing sparsification. Routing exhibits a clear threshold between removing the lowest-weight route for each token and retaining only two routes. Keeping the three largest realized read or write weights for every token at every site raises perplexity by only 2.4–2.7% and changes the average downstream score by at most 0.38 percentage points. Keeping only the two largest weights per token, however, raises perplexity by 12.1–14.0% and reduces the average downstream score by 2.92–6.59 points; for 𝑯 post , the larger average drop is driven primarily by GSM8K. The lowestweight routing path for each token is therefore largely dispensable under the evaluated interventions. Consistent with this redundancy being local rather than stream-wide, Appendix A.2 (Figure 8) shows that the stream with the lowest mean routing load changes across layers, maps, and sublayer types. Residual-mixer replacement. The functional importance of learned residual mixing closely follows its depth profile. Replacing the near-identity mixers in layers 22–42 by exact identity maps raises perplexity by only 1.9% and changes the six-task average score by +0.04 percentage points. In contrast, applying the same replacement to layers 0–21 raises perplexity by 41.4%, nearly matching the 42.3% increase from replacing the mixers throughout the network. Fixing each early mixer to its site-specific mean over the C4 diagnostic sequences yields a perplexity of 11.1832, only 0.2% above the 11.1587 baseline, and changes the six-task average score from 84.22% to 83.98% (−0.25 percentage points). For this checkpoint and the evaluated metrics, the learned site-specific structure therefore accounts for most of the measured benefit of the early mixers, while their token-wise vari9

Preprint.

ation contributes little; the late mixers can instead be replaced by identity with little measured loss, eliminating their Sinkhorn–Knopp iterations at inference time. Together, these interventions show that training realizes only part of the flexibility afforded by mHC’s four-stream design, with redundancy localized both within routing maps and across depth in the residual mixers. This is not a wholesale collapse of multi-stream computation: the remaining routing support and the site-specific structure of the early mixers materially affect model quality.

5

D ISCUSSION AND I MPLICATIONS

Our findings identify two properties of the trained pathway that call for further analysis: routing is concentrated within individual sublayers but reorganizes across depth, and learned residual mixing becomes nearly identity after mid-depth. We identify candidate optimization mechanisms associated with these patterns and consider implications for future multi-stream designs. A candidate feedback mechanism for routing concentration. Read/write routing and stream representations are learned jointly, creating a possible feedback loop: greater read weight propagates a larger branch-mediated gradient to the parameters producing that stream, while the write map determines which persistent streams receive the branch output. These interactions may reinforce existing routing preferences, although the realized coefficients remain coupled through a shared token-dependent router. Appendix C.1 formalizes this channel; validating it requires training-time trajectories of routing margins, stream-wise gradients, and representations. Optimization factors governing residual mixing. Movement away from identity depends on both the optimization pressure for cross-stream exchange and the sensitivity of Sinkhorn–Knopp normalization to that pressure. The final checkpoint cannot distinguish weak functional demand from weak gradient transmission to off-diagonal entries; Appendix C.2 formalizes this distinction. Neither local analysis reconstructs the training dynamics of the checkpoint, whose causal mechanisms require training-time routing and mixer trajectories. Directions for improving mHC. Recent designs suggest that multi-stream residual parameterization remains unsettled: Hy4-preview adopts four-stream identity Hyper-Connections (Tencent Hy Team, 2026), while Qwen3.8-Flash-Next uses data-dependent Gated Residual across four branches (Qwen Team, 2026). These choices accord with our finding that not every degree of freedom in full mHC is functionally required. In mHC, all three maps are predicted from incoming states: this is necessary for 𝑯 pre , but 𝑯 post selects where to write the branch output and 𝑯 res mixes incoming states without conditioning on the computation produced by the branch. A promising alternative is to retain input-conditioned 𝑯 pre while allowing 𝑯 post , and potentially 𝑯 res , to condition additionally on the branch output; evaluating it requires matched training comparisons with full and simplified alternatives.

6

C ONCLUSION

We studied how a trained mHC model realizes the flexibility of its multi-stream residual pathway during inference. True-forward measurements of DeepSeek-V4-Flash reveal locally concentrated but depth-varying routing, directionally distinct stream representations, and residual exchange that occurs primarily in early layers. Inference-time interventions further show that the weakest routing weight per token at every site and the late residual mixers can be suppressed with little measured loss, whereas the remaining routing support and the learned structure of the early mixers are functionally important. Fixing the early mixers to their C4 diagnostic means also nearly preserves model quality, indicating that their site-specific structure contributes more than their token-wise variation on the evaluated metrics. The four-stream pathway thus exhibits structured under-utilization rather than uniform participation or complete stream collapse. More broadly, architectural width does not determine effective multistream capacity; its realization depends on the component- and depth-specific organization learned during training. Our analysis identifies candidate mechanisms whose causal roles require trainingtime routing and mixer trajectories. 10

Preprint.

R EFERENCES Ekaterina Alimaskina, Gleb Molodtsov, and Aleksandr Beznosikov. Analyzing stream collapse in Hyper-Connections: From diagnosis to mitigation, 2026. URL https://arxiv.org/abs/ 2606.03483. Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: Reasoning about physical commonsense in natural language. In AAAI, 2020. URL https://arxiv. org/abs/1911.11641. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Øyvind Tafjord. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. URL https://arxiv.org/abs/1803. 05457. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. URL https://arxiv.org/abs/2110.14168. DeepSeek-AI. DeepSeek-V4: Towards highly efficient Million-Token context intelligence, 2026. URL https://arxiv.org/abs/2606.19348. Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: Accurate post-training quantization for generative pre-trained transformers. In International Conference on Learning Representations, 2023. URL https://arxiv.org/abs/2210.17323. Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, 2023. URL https://zenodo.org/records/ 10256836. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In ICLR, 2021. URL https: //arxiv.org/abs/2009.03300. Andrej Karpathy. nanoGPT: The simplest, fastest repository for training and fine-tuning mediumsized GPTs. https://github.com/karpathy/nanoGPT, 2022. GitHub repository. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, volume 32, 2019. Qwen Team. On the design of Qwen3.8-Next architecture: Evaluation, efficiency, and training stability. Technical report, Alibaba Group, August 2026. URL https://arxiv.org/abs/ 2608.30320. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(140):1–67, 2020. Richard Sinkhorn and Paul Knopp. Concerning nonnegative matrices and doubly stochastic matrices. Pacific Journal of Mathematics, 21(2):343–348, 1967. Tencent Hy Team. Hy4-preview model card. https://huggingface.co/tencent/ Hy4-preview, 2026. Accessed September 2026. 11

Preprint.

Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Perric Cistac, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, 2020. Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Jiang Chang, Kuai Yu, Liang Zhao, Shangyan Zhou, Zhean Xu, Zhengyan Zhang, Wangding Zeng, Shengding Hu, Yuqing Wang, Jingyang Yuan, Lean Wang, and Wenfeng Liang. mHC: Manifold-Constrained Hyper-Connections. arXiv preprint arXiv:2512.24880, 2025. URL https://arxiv.org/abs/2512.24880. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800, 2019. doi: 10.18653/v1/P19-1472. URL https://aclanthology.org/P19-1472/. Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, and Xun Zhou. Hyper-Connections. arXiv preprint arXiv:2409.19606, 2024. URL https: //arxiv.org/abs/2409.19606.

12

Preprint.

A

A DDITIONAL D IAGNOSTIC R ESULTS

A.1

D IAGNOSTICS BY D EPTH AND S UBLAYER T YPE

Table 3 places the routing, representation, and residual-mixing diagnostics under a common depth split. The reduction in residual mixing is consistent across attention and FFN sites, while the change in routing breadth is concentrated in the write maps, particularly at FFN sites. Table 3: Mean diagnostics by depth range and sublayer type. Winner consistency (WC) and effective stream count (𝑛eff ) are averaged over the indicated routing sites; cosine similarity is additionally averaged over the six stream pairs. Winner consistency

𝑛eff

Cosine

𝛿(𝑯 res )

Layers

Site

𝑯 pre

𝑯 post

𝑯 pre

𝑯 post

0–21

Attention FFN

0.874 0.905

0.970 0.909

2.174 1.768

1.317 1.618

0.456 0.438

0.045 0.047

22–42

Attention FFN

0.833 0.872

0.917 0.823

2.104 1.947

1.836 2.360

0.385 0.378

0.010 0.008

A.2

I DENTITY OF THE L OWEST-L OAD ROUTE

Figure 8 ranks streams by their mean normalized routing load at each site. All four streams appear as the lowest-load stream somewhere in each map/site combination, so the locally weak route does not correspond to a single stream that can be removed throughout the network. Its mean load is generally small but varies across depth, favoring site-adaptive sparsification over the removal of a globally fixed stream. (a)

Stream 0

Stream 1

Stream 2

Stream 3

H pre (Attn) H pre (FFN) H post (Attn) H post (FFN)

(b)

0.20

H pre (FFN)

0.15

H post (Attn) H

post

0.10 0.05

(FFN)

Lowest load

0.25

H pre (Attn)

0.00 0

5

10

15

20

25

30

35

40

42

Layer

Figure 8: Lowest-load routing stream across depth. (a) Identity of the stream with the smallest mean normalized routing load at each layer, map, and sublayer type. (b) Its corresponding mean load. The weakest stream changes across sites rather than remaining globally fixed.

13

Preprint.

B

L AYERWISE R ESIDUAL M IXERS

Figures 9 and 10 show the token-averaged realized residual mixer at every layer, separately for attention and FFN sites. They complement the depth profile and range averages in Figure 7 by exposing which stream pairs account for the residual exchange at each site. Attention sites. Figure 9 shows that the off-diagonal structure is both depth- and pair-specific. The most visible exchange occurs at several early and middle sites, including layers 1–2, 9, and 11, followed by a final pronounced deviation at layer 22. Beyond this point, the matrices are consistently near diagonal, with only small site-level fluctuations through layer 42. L00

L01

L02

L03

L04

L05

0.92 0.00 0.00 0.00

0.83 0.00 0.13 0.03

0.60 0.00 0.39 0.00

0.87 0.00 0.10 0.03

0.94 0.01 0.00 0.03

0.92 0.07 0.00 0.00

0.00 0.88 0.06 0.10

0.17 0.70 0.14 0.00

0.05 1.00 0.00 0.00

0.02 1.00 0.03 0.00

0.00 0.78 0.24 0.00

0.08 0.79 0.10 0.03

0.00 0.11 0.94 0.01

0.00 0.30 0.72 0.00

0.35 0.00 0.61 0.02

0.11 0.00 0.87 0.02

0.05 0.20 0.76 0.01

0.00 0.12 0.90 0.00

0.08 0.00 0.00 0.89

0.00 0.00 0.00 0.97

0.00 0.00 0.00 0.98

0.00 0.00 0.00 0.95

0.00 0.00 0.00 0.96

0.00 0.02 0.00 0.97

L06

L07

L08

L09

L10

L11

0.95 0.02 0.00 0.01

0.97 0.01 0.00 0.00

0.95 0.05 0.01 0.00

0.93 0.05 0.03 0.00

1.00 0.00 0.03 0.00

0.82 0.10 0.08 0.01

0.04 0.83 0.12 0.03

0.02 0.94 0.00 0.04

0.01 0.93 0.00 0.04

0.00 0.63 0.36 0.01

0.00 0.94 0.03 0.00

0.03 0.71 0.25 0.01

0.01 0.14 0.88 0.00

0.00 0.04 0.99 0.00

0.04 0.01 0.99 0.00

0.07 0.32 0.59 0.01

0.00 0.06 0.93 0.01

0.15 0.19 0.65 0.00

0.00 0.01 0.00 0.97

0.00 0.01 0.00 0.96

0.00 0.01 0.00 0.95

0.00 0.00 0.02 0.97

0.00 0.00 0.01 0.99

0.00 0.00 0.02 0.98

L12

L13

L14

L15

L16

L17

0.97 0.00 0.05 0.00

0.99 0.05 0.00 0.00

0.99 0.05 0.00 0.00

1.00 0.04 0.00 0.01

0.99 0.03 0.00 0.02

1.00 0.01 0.00 0.03

0.00 0.86 0.11 0.00

0.01 0.92 0.06 0.00

0.00 0.89 0.09 0.00

0.00 0.81 0.16 0.00

0.00 0.86 0.12 0.00

0.00 0.91 0.06 0.00

0.02 0.14 0.82 0.00

0.00 0.01 0.94 0.00

0.00 0.04 0.91 0.00

0.00 0.12 0.84 0.00

0.00 0.07 0.88 0.00

0.00 0.06 0.91 0.00

0.00 0.00 0.02 1.00

0.00 0.02 0.00 1.00

0.00 0.03 0.00 1.00

0.00 0.03 0.00 0.99

0.00 0.04 0.00 0.98

0.00 0.02 0.03 0.97

L18

L19

L20

L21

L22

1.00 0.00 0.03 0.01

0.99 0.00 0.03 0.01

0.97 0.00 0.04 0.02

0.98 0.01 0.04 0.00

0.77 0.02 0.21 0.00

0.00 0.94 0.01 0.00

0.00 0.94 0.00 0.00

0.00 0.95 0.00 0.00

0.00 0.95 0.00 0.00

0.00 0.96 0.00 0.00

0.00 0.06 0.92 0.00

0.00 0.06 0.92 0.00

0.02 0.03 0.94 0.01

0.00 0.04 0.95 0.00

0.20 0.02 0.78 0.00

0.00 0.00 0.04 0.99

0.00 0.00 0.04 0.99

0.01 0.01 0.01 0.98

0.02 0.00 0.01 1.00

0.03 0.00 0.01 1.00

L23

L24

L25

L26

L27

0.93 0.04 0.03 0.00

0.96 0.02 0.01 0.00

0.99 0.00 0.00 0.00

0.99 0.01 0.00 0.00

0.99 0.01 0.01 0.00

0.00 0.95 0.00 0.00

0.00 0.97 0.00 0.00

0.00 0.99 0.01 0.00

0.00 0.98 0.00 0.00

0.00 0.99 0.00 0.00

0.06 0.01 0.94 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.99 0.00

0.00 0.00 0.98 0.00

0.01 0.00 0.03 1.00

0.03 0.00 0.00 1.00

0.01 0.01 0.01 1.00

0.01 0.01 0.01 1.00

0.01 0.00 0.01 1.00

L28

L29

L30

L31

L32

0.99 0.00 0.00 0.00

0.99 0.00 0.00 0.00

1.00 0.00 0.00 0.00

1.00 0.00 0.00 0.00

0.99 0.00 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.98 0.00

0.00 0.01 0.98 0.00

0.00 0.01 0.98 0.00

0.00 0.00 0.98 0.00

0.00 0.01 0.02 1.00

0.01 0.00 0.01 1.00

0.00 0.01 0.02 1.00

0.00 0.00 0.02 1.00

0.00 0.00 0.01 1.00

L33

L34

L35

L36

L37

0.99 0.01 0.00 0.00

0.99 0.01 0.00 0.00

0.98 0.01 0.01 0.00

0.99 0.01 0.00 0.00

0.99 0.01 0.01 0.00

0.00 0.99 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.99 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.99 0.00

0.01 0.00 0.01 1.00

0.01 0.01 0.01 1.00

0.01 0.00 0.01 1.00

0.01 0.01 0.01 0.99

0.01 0.00 0.00 1.00

L38

L39

L40

L41

L42

0.99 0.00 0.01 0.01

0.99 0.00 0.01 0.00

0.95 0.01 0.03 0.01

0.98 0.00 0.01 0.01

0.97 0.01 0.01 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.01 0.98 0.00 0.00

0.01 1.00 0.01 0.00

0.01 0.98 0.01 0.01

0.00 0.00 0.98 0.01

0.01 0.00 0.98 0.00

0.02 0.00 0.96 0.00

0.01 0.00 0.96 0.02

0.01 0.01 0.97 0.02

0.00 0.01 0.01 0.98

0.01 0.01 0.01 0.99

0.03 0.00 0.00 0.98

0.01 0.00 0.03 0.97

0.01 0.01 0.02 0.96

Figure 9: Token-averaged realized residual mixers at the attention site of every layer.

14

1.0

0.8

0.6

0.4

0.2

0.0

Preprint.

FFN sites. Figure 10 exhibits the same broad reduction in residual exchange but a different set of high-mixing sites. The strongest off-diagonal weights occur at layers 0–1 and 13, whereas the mixers from layer 23 onward remain close to identity. The locations and stream pairs responsible for early mixing therefore differ between attention and FFN, while the late near-identity regime is shared by both sublayer types. L00

L01

L02

L03

L04

L05

0.26 0.00 0.00 0.71

0.24 0.00 0.73 0.02

0.88 0.00 0.10 0.01

0.92 0.01 0.03 0.02

0.97 0.00 0.00 0.00

0.95 0.01 0.00 0.01

0.00 0.98 0.06 0.00

0.04 1.00 0.01 0.00

0.05 1.00 0.00 0.00

0.01 0.84 0.19 0.00

0.00 0.95 0.07 0.01

0.04 0.91 0.03 0.03

0.01 0.00 0.94 0.06

0.71 0.00 0.26 0.02

0.07 0.00 0.90 0.01

0.07 0.15 0.78 0.02

0.03 0.05 0.93 0.01

0.00 0.07 0.97 0.00

0.73 0.02 0.00 0.23

0.01 0.00 0.00 0.96

0.00 0.00 0.00 0.97

0.00 0.00 0.00 0.96

0.00 0.00 0.00 0.98

0.00 0.01 0.00 0.97

L06

L07

L08

L09

L10

L11

0.92 0.04 0.00 0.00

0.94 0.03 0.00 0.01

0.95 0.04 0.00 0.01

1.00 0.00 0.03 0.00

0.99 0.00 0.04 0.00

0.99 0.00 0.04 0.00

0.08 0.89 0.02 0.01

0.06 0.91 0.01 0.01

0.00 0.95 0.00 0.00

0.00 0.94 0.02 0.00

0.00 0.94 0.02 0.00

0.00 0.91 0.04 0.00

0.00 0.05 0.98 0.00

0.01 0.04 0.99 0.00

0.05 0.00 0.99 0.00

0.00 0.06 0.93 0.00

0.01 0.05 0.93 0.00

0.01 0.09 0.89 0.00

0.00 0.02 0.00 0.99

0.00 0.01 0.00 0.98

0.00 0.01 0.00 0.99

0.00 0.00 0.02 1.00

0.00 0.00 0.01 1.00

0.00 0.00 0.02 1.00

L12

L13

L14

L15

L16

L17

1.00 0.02 0.02 0.00

0.99 0.04 0.00 0.00

0.99 0.05 0.00 0.00

0.99 0.05 0.00 0.01

0.99 0.04 0.00 0.01

1.00 0.00 0.04 0.00

0.00 0.93 0.05 0.00

0.01 0.28 0.69 0.00

0.00 0.85 0.12 0.00

0.00 0.90 0.08 0.00

0.00 0.91 0.07 0.00

0.00 0.94 0.00 0.00

0.00 0.05 0.92 0.00

0.00 0.66 0.31 0.00

0.01 0.07 0.88 0.00

0.00 0.02 0.92 0.00

0.00 0.03 0.92 0.00

0.00 0.05 0.93 0.00

0.00 0.00 0.01 1.00

0.00 0.02 0.00 1.00

0.01 0.03 0.00 1.00

0.01 0.03 0.00 0.99

0.01 0.03 0.00 0.99

0.00 0.00 0.02 1.00

L18

L19

L20

L21

L22

1.00 0.00 0.03 0.01

0.99 0.00 0.05 0.00

0.98 0.00 0.05 0.00

0.97 0.01 0.05 0.00

0.95 0.03 0.00 0.00

0.00 0.94 0.00 0.00

0.00 0.94 0.00 0.00

0.00 0.95 0.00 0.00

0.00 0.96 0.00 0.00

0.00 0.95 0.00 0.00

0.00 0.06 0.93 0.00

0.00 0.06 0.93 0.00

0.00 0.05 0.94 0.00

0.01 0.04 0.95 0.00

0.03 0.01 0.98 0.00

0.00 0.00 0.04 0.98

0.01 0.00 0.02 1.00

0.01 0.00 0.01 1.00

0.02 0.00 0.00 1.00

0.02 0.00 0.01 1.00

L23

L24

L25

L26

L27

0.97 0.03 0.01 0.00

0.98 0.00 0.00 0.00

0.97 0.00 0.00 0.00

0.98 0.02 0.00 0.00

0.98 0.00 0.00 0.00

0.00 0.97 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.00 0.99 0.00

0.01 0.01 0.98 0.00

0.02 0.00 0.98 0.00

0.01 0.00 0.98 0.00

0.02 0.00 0.98 0.00

0.03 0.00 0.01 1.00

0.01 0.01 0.02 1.00

0.01 0.01 0.02 1.00

0.01 0.00 0.02 1.00

0.00 0.01 0.02 1.00

L28

L29

L30

L31

L32

1.00 0.00 0.00 0.00

1.00 0.00 0.00 0.00

0.99 0.01 0.00 0.00

0.99 0.00 0.00 0.00

0.99 0.00 0.00 0.00

0.00 0.99 0.01 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.99 0.00

0.00 0.00 0.98 0.00

0.01 0.02 0.97 0.00

0.01 0.01 0.98 0.00

0.00 0.01 0.02 1.00

0.00 0.01 0.01 1.00

0.01 0.00 0.01 1.00

0.00 0.01 0.03 1.00

0.00 0.00 0.02 1.00

L33

L34

L35

L36

L37

0.99 0.02 0.00 0.00

0.99 0.01 0.00 0.00

0.97 0.03 0.01 0.00

0.99 0.01 0.00 0.00

0.98 0.01 0.01 0.01

0.00 0.97 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.97 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.98 0.00 0.00

0.00 0.01 0.98 0.00

0.00 0.00 0.99 0.00

0.00 0.00 0.99 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.98 0.00

0.01 0.01 0.01 1.00

0.01 0.01 0.01 1.00

0.03 0.00 0.00 1.00

0.00 0.00 0.01 0.99

0.02 0.01 0.01 0.99

L38

L39

L40

L41

L42

0.99 0.01 0.01 0.00

0.98 0.00 0.01 0.01

0.96 0.00 0.04 0.00

0.99 0.00 0.01 0.00

0.99 0.01 0.01 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.99 0.00 0.00

0.00 0.97 0.00 0.00

0.00 0.00 0.98 0.00

0.00 0.00 0.97 0.01

0.03 0.00 0.95 0.01

0.00 0.00 0.98 0.00

0.00 0.01 0.98 0.01

0.01 0.00 0.01 1.00

0.01 0.00 0.02 0.98

0.01 0.00 0.01 0.99

0.01 0.00 0.01 0.99

0.01 0.01 0.01 0.99

Figure 10: Token-averaged realized residual mixers at the FFN site of every layer.

15

1.0

0.8

0.6

0.4

0.2

0.0

Preprint.

C

L OCAL O PTIMIZATION A NALYSIS

C.1

ROUTING C ONCENTRATION

For a single token and sublayer, Í write the realized read and write maps as 𝒓 = (𝑟 1 , . . . , 𝑟 𝑛 ) and 𝒘 = (𝑤 1 , . . . , 𝑤 𝑛 ), such that 𝒛 = 𝑗 𝑟 𝑗 𝒙 ( 𝑗 ) , 𝒚 = 𝐹 (𝒛), and the branch adds 𝑤 𝑖 𝒚 to output stream 𝑖. Let 𝒈 𝑧 = 𝜕L/𝜕𝒛 and 𝒈𝑖+ = 𝜕L/𝜕𝒙 +(𝑖) . Conditioned on the realized routing coefficients, the direct branch-mediated gradients are 𝜕L = 𝑟 𝑗 𝒈𝑧 , 𝜕𝒙 ( 𝑗 ) branch

𝜕L = 𝒈𝑧 , 𝒙 ( 𝑗 ) , 𝜕𝑟 𝑗

𝜕L = 𝒈𝑖+ , 𝒚 . 𝜕𝑤 𝑖

(12)

At the level of the realized read coefficients, define the negative-gradient direction 𝑑𝑟 𝑗 := −

𝜕L = − 𝒈𝑧 , 𝒙 ( 𝑗 ) . 𝜕𝑟 𝑗

(13)

Stream 𝑗 is favored over stream 𝑘 by this local direction when 𝑑𝑟 𝑗 > 𝑑𝑟𝑘 . A larger 𝑟 𝑗 also scales the branch-mediated gradient propagated through 𝒙 ( 𝑗 ) , and hence the gradient received by the upstream parameters that produce this state. If the resulting parameter update makes 𝑑𝑟 𝑗 more favorable on subsequent examples, the existing read preference can be reinforced. The write map participates in the same feedback by controlling which persistent states receive the branch output and influence later computation. Because the realized routing coefficients are coupled through a shared tokendependent router, these descent scores identify a possible feedback channel rather than the router parameters’ actual update. C.2

R ESIDUAL M IXING

Write the realized mixer at a token and sublayer as 𝑯 res = S( 𝑨), where 𝑨 contains the pre-Sinkhorn logits and S denotes Sinkhorn normalization, and let 𝒉 := vec(𝑯 res ),

𝒂 := vec( 𝑨),

𝑱 S :=

𝜕𝒉 , 𝜕𝒂

𝒈 ℎ :=

𝜕L . 𝜕𝒉

(14)

For an idealized local gradient step on the logits, first-order expansion gives Δ𝒉 ≈ −𝜂 𝑱 S 𝑱 ⊤ S 𝒈ℎ .

(15)

If 𝑷off projects a vectorized mixer onto its off-diagonal entries, then ∥𝑷off Δ𝒉∥ 2 ≤ 𝜂 ∥𝑷off 𝑱 S ∥ 2 ∥ 𝑱 S ∥ 2 ∥ 𝒈 ℎ ∥ 2 .

(16)

Movement away from identity can therefore remain small either because the loss gradient supplies little pressure for cross-stream residual exchange or because the local Sinkhorn Jacobian weakly transmits that pressure to off-diagonal entries. The final checkpoint does not reveal which factor governed the late-layer mixers or how they reached the observed regime.

16

Record · ID 660800 · SHA-256 0d9a8332380d8f4e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.