Conceptio › Archive › arXiv CS
arXiv CSopen access

Affinity-Aware Sharding for Delayed Tensor Parallelism

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Affinity-Aware Sharding for Delayed Tensor Parallelism Eloi de Reynal [email protected]

arXiv:2609.13846v1 [cs.LG] 12 Sep 2026

Abstract Delayed Tensor Parallelism (DTP) removes the blocking all-reduce of tensor-parallel Transformer inference. Every device adds its own partial output to its residual stream (and broadcasts it) immediately, but only gathers (receives) the other devices’ partials δ modules later. This architecture forces some degree of independence between devices, as the communication delay also degrades its quality: each partial is computed from a stream that lacks the other devices’ latest writes. A TP to DTP change therefore amounts to a real architecture change, and dense Transformer models need to be retrained or distilled after adaptation. We show that DTP breaks the permutation symmetry of neurons inside FFNs and of KV heads inside attention modules, and that this symmetry breakage makes the sharding itself a modelling decision. We show that maximising the affinity between the KV heads and the FFN neurons co-located on a device, by permuting the dense model before sharding, speeds up the distillation or retraining process. The affinity is measured with a first-order approximation of the damage that losing a head’s contribution does to each neuron’s output, and the co-located affinity is maximised with a coordinate-ascent optimiser that alternates an exact balanced assignment of neurons with an exhaustive search over the KV head partitions. The whole procedure takes under two minutes on one GPU for Qwen3-0.6B and Danube3-500M. On these models, at δ = 1, the affinity-optimised layouts reach any distillation target in about half to two thirds of the steps needed by the naive contiguous layouts, over the whole 10k-step range we tested, and every optimised seed beats every contiguous seed and all but one of the sixteen random layouts. We also show that the co-located affinity score at initialisation predicts the KL to the base model after training, across seventeen layouts ranging from anti-optimised to optimised (Pearson −0.81 and −0.89). Ablations show that the gain comes mostly from placing neurons (more than heads), and that zero-shot damage (as opposed to zero-shot affinity) is a poor predictor of the trained quality.

1

Introduction

Batch-size-one decoding of a Transformer is bound by weight streaming, and can therefore be sped up with tensor parallelism (TP), which divides the amount of weights to be streamed on each device by the total number of devices (Shoeybi et al., 2019). Communication overhead soon becomes an issue at low batch sizes, as the computations themselves become fast and shorter than the communication. At low batch sizes again, the communication speed is limited by incompressible latency more than by bandwidth itself. As a consequence, the all-reduce becomes a large share of the per-token latency: removing it improves the decode throughput of Llama-3.1-8B on eight H200s by 43% (Taneja et al., 2026). Various works focus on changing the architecture to hide communication behind compute: • parallel attention and MLP blocks halve the number of synchronisations, at the cost of also halving the critical-path depth (Chowdhery et al., 2023; Wang and Komatsuzaki, 2021); • Ladder Residual feeds each module the previous module’s residual, so that compute happens while the all-reduce finishes (Zhang et al., 2025); • Kraken and Parallel Track Transformers run several thinner Transformers side by side with periodic syncs, at the cost of lowering the effective width of each layer (Prabhakar et al., 2024; Wang et al., 2026); • Sync-Point Drop skips the attention all-reduce in insensitive blocks, and therefore sometimes all-reduces only after the FFN (Kim et al., 2025);

1

• CAAT-Net synchronises only a subset of channels (Lamprecht et al., 2025). Delayed Tensor Parallelism (DTP) (Kog Team, 2026) is the most recent addition, and our paper builds on it. In DTP, every device adds its own partial output to its residual stream (and broadcasts it) immediately, but only gathers (receives) the other devices’ partials δ modules later. At δ = 1, and from the second module up to the penultimate one, the input of each module is the full residual stream minus the other devices’ contribution to the immediately preceding module. In other words, the input of module n on device d is the residual stream of device d after module n − 1, plus the contribution of all other devices after module n − 2. As DTP changes the function the network computes, a pretrained dense model has to be retrained or distilled after adaptation, like every other architecture mentioned above (Kim et al., 2025; Zhang et al., 2025). DTP also breaks the permutation symmetry of neurons inside FFNs and of KV heads inside attention modules, and therefore makes the sharding a non-neutral choice that can be optimised. Under DTP, the FFN of layer n on device d only sees the attention output of its co-located set of preceding KV heads. If the neurons whose output depends most on these KV heads indeed sit on the same device, little information is lost. If, on the contrary, the neurons most sensitive to these KV heads’ output lie on other devices, most of the information they output will be lost for layer n. Since, for the original dense model, the heads of an attention layer and the neurons of an FFN are permutation-symmetric (Ainsworth et al., 2023; Entezari et al., 2022; Tran et al., 2026), we are free to choose any balanced assignment of heads and neurons to devices. The permutation only changes the DTP function, while leaving the dense model unchanged to within floating-point error. Contributions. 1. We formalise the choice of sharding under DTP as a layout optimisation problem over permutations of key-value (KV) groups and FFN neurons, and give a calibration-based affinity score between heads and neurons (a first-order estimate of the damage to the residual stream when a neuron loses one head’s contribution) and between neurons and the next layer’s KV groups (§3). 2. We give a fast optimiser for the resulting chain of assignment problems: an exact balanced linear assignment for the neurons of a layer given its neighbouring head partitions, and an exhaustive search over the 2520 labelled balanced partitions of eight KV groups into four devices given the neighbouring neuron assignments, alternated until convergence. Affinity pass plus optimisation take under two minutes on one GPU. 3. On two pretrained models sharded over four devices at δ = 1, distilled from the dense model with an identical recipe for every layout, we show that (a) every optimised seed beats every contiguous seed and every random layout on Qwen3-0.6B (seven of eight on Danube3-500M), and the optimised layout reaches the contiguous run’s 1000-step quality in 347 steps on Qwen3-0.6B and 565 on Danube3-500M; (b) over 10k steps the optimised run reaches each contiguous KL value in about half to two thirds of the steps, with no sign of the gap closing on Qwen3-0.6B; (c) the co-located affinity of a layout predicts its trained KL across seventeen layouts; (d) the gain lives in the neuron assignment, the head assignment adds nothing measurable, and the zero-shot damage of a layout is uninformative about its trained quality (§5).

2

Background: Delayed Tensor Parallelism at δ = 1

We consider a pre-norm decoder with N layers, each an attention module followed by an FFN module, indexed jointly by n = 0, . . . , 2N − 1 (attention of layer i is module 2i, its FFN module 2i + 1). Sharded over L devices in the Megatron way (Shoeybi et al., 2019), attention is split by heads (by KV groups under grouped-query attention (Ainslie et al., 2023), each group carrying its query heads) and the FFN by intermediate P neurons, so that module n on device l produces a partial output onl and the vanilla module output is l onl . Standard TP all-reduces after every module; each device then continues from the same residual stream. DTP gives each device its own residual stream Xl and delays the aggregation by δ modules (Kog

2

S (2i−1) : every device’s partials up to the FFN of layer i − 1

attention i

heads on device j ′ ̸= l

FFN of layer i, the neurons on device l

FFN i

attention i + 1

heads on device j ̸= l

heads on device l

KV groups on device j ̸= l

KV groups on device l

KV groups on device j ′ ̸= l

dashed: the other devices’ partial of the preceding module, which lands one module later

Figure 1: What device l sees under DTP at δ = 1 (equation 4). Solid: available when the module runs. Dashed red: not yet available. The two edges the layout can control are heads(i) →neurons(i) (the up edge) and neurons(i) →KV groups(i + 1) (the down edge).

Team, 2026): n<δ: δ ≤ n < 2N − δ : n ≥ 2N − δ :

√

L onl , P Xl ← Xl + onl + j̸=l ojn−δ , √ P Xl ← Xl + L onl + j̸=l ojn−δ ,

Xl ← Xl +

(1) (2) (3)

√ where the L factor in the first and last δ modules mimics the magnitude of a full all-reduce. At the end the final norm is applied per device and the L outputs are averaged. The broadcast of onl has δ modules of weight streaming to arrive, which is what removes the exposed communication. This paper studies δ = 1, the smallest delay and the one where the structure is cleanest. Unrolling (2) at δ = 1 shows that before module n + 1 device l holds P P (n) Xl = S (n−1) + onl , S (m) = X (0) + k≤m j okj , (4) that is, the sum of every device’s partials up to module n − 1, plus only its own partial of module n (Figure 1). Two things follow. The FFN of layer i on device l sees the attention of layer i only through the KV groups placed on l; the attention of layer i + 1 on device l sees the FFN of layer i only through the neurons placed on l. Everything older is complete. The partials themselves are computed from diverged streams, so the model is not the dense model and must be trained; but what a device misses is determined entirely by which heads and neurons share the device.

3

Method: affinity-aware layouts

The layout as a free variable. For layer i let hi : {1..G} → {1..L} assign its G KV groups to devices and νi : {1..I} → {1..L} assign its I FFN neurons, both balanced (G/L groups and I/L neurons per device). Any such pair is realised by permuting the rows of the query, key and value projections (by group, keeping each group’s query heads with its KV heads) and the columns of the output projection identically, and the rows of the gate and up projections and the columns of the down projection identically, then slicing contiguous shards. The dense function is unchanged: we verify a maximum logit difference of 10−4 in fp32. The rotary embedding, the per-head norms of Qwen3 and the SwiGLU nonlinearity (Shazeer, 2020) are all applied per head or per neuron, so they commute with the permutation. Affinity scores. Per layer we need two matrices: Si↑ ∈ RG×I , how much neuron k of layer i depends on KV group g of layer i, and Si↓ ∈ RI×G , how much KV group g of layer i + 1 depends on neuron k of layer i. Both are estimated from the dense model on calibration text (64 sequences of 1024 tokens of FineWeb-Edu (Penedo et al., 2024), one forward pass, no gradients). We write x for the vanilla residual entering the FFN’s pre-norm, r = rms(x), γ the norm gain, ah head h’s attention 3

(h)

output and Wo its slice of the output projection, so that the normalised FFN input decomposes P (h) additively over heads as h ph + (rest) with ph P = γ ⊙ (ah Wo )/r (ElhageP et al., 2021). Neuron g k’s pre-activations therefore decompose as gk = h ph ·wk + . . . and uk = h ph ·wku + . . . , and the first-order change of its output silu(gk ) uk when head h’s piece is removed is ∆hk = silu′ (gk ) uk (ph ·wkg ) + silu(gk ) (ph ·wku ).

(5)

Our main score, FO (first order), is the mean squared damage this does to the residual stream, ↑,FO Shk = Et [∆2hk ] ∥wkd ∥2 with wkd the down-projection column, summed over the query heads of each KV group. The simpler ADD score is the mean squared additive contribution to the pre-activations, Et [(ph · wkg )2 + (ph · wku )2 ]; it ignores the operating point of the nonlinearity and the size of the neuron’s write. For the down edge, neuron k writes αk wkd into the residual (αk its activation), which enters the next layer’s q/k/v projections of group g through that layer’s pre-norm (gain γ ′ , rms r′ ) as ↓ bgk = Wgqkv (γ ′ ⊙ wkd ), so Skg = Et [αk2 /r′2 ] ∥bgk ∥2 . We also computed exact leave-one-out ablation scores of the pre-activations, which differ from ADD only by including the change of the norm’s rms when a head is removed; their Spearman correlation with ADD is above 0.98 on every layer but one (0.83 on the last layer of Danube3-500M), so the rms change is negligible and both ADD and FO ignore it. Objective. Each matrix is normalised to unit mass so that every edge counts equally, and the layout score is the total co-located affinity, J(h, ν) =

N −1 X

X

Ŝi↑ [g, k] +

i=0 g,k: hi (g)=νi (k)

N −2 X

X

Ŝi↓ [k, g],

(6)

i=0 k,g: νi (k)=hi+1 (g)

which reads as “number of edges times mean co-located fraction”: it ranges over [0, 2N − 1] and a random balanced layout scores (2N − 1)/L in expectation. J is a proxy, not the training loss; §5.2 tests how well it predicts the trained outcome. Optimiser. J couples the layers in a chain (Figure 1), and each block of the chain is easy given its neighbours. Given hi and hi+1 , the gain of putting neuron k on device l is X X Mkl = Ŝi↑ [g, k] + Ŝi↓ [k, g], g: hi (g)=l

g: hi+1 (g)=l

and the best balanced νi is a linear assignment problem on I/L replicated slots per device (Crouse, 2016; Kuhn, 1955). We solve it on the GPU by dual ascent on L device offsets, a fix-up to meet the capacities exactly, and pairwise swaps until no swap improves; on the layers we checked it matches the Hungarian solution to 10−6 . Given νi−1 and νi , the best hi is found by scoring all G!/((G/L)!)L labelled balanced partitions of the KV groups (2520 for G = 8, L = 4). Starting from the contiguous head assignment we alternate the two steps over all layers until J stops increasing, which takes two to four sweeps. Restricting which layers move, freezing heads or neurons, or minimising J instead gives the ablation layouts of §5.4. Cost. The affinity pass takes 49 s (Danube3-500M) and 51 s (Qwen3-0.6B) on one RTX 5090 for 65,536 tokens; the optimisation 38 to 47 s including scoring eight random layouts. One 1000-step distillation run takes 22 and 32 min on the same GPU, so the method costs about 7% of the shortest run in this study.

4

Experimental setup

Models and sharding. Qwen3-0.6B (Qwen Team, 2025) (28 layers, 16 query and 8 KV heads, head dim 128, FFN width 3072) and Danube3-500M (Pfeiffer et al., 2024) (16 layers, 16 query and 8 KV heads, FFN 4096), both pre-norm SwiGLU decoders with the Llama module layout. L = 4 virtual devices are simulated√ on one GPU (the sharding arithmetic is exact and tested; the speed-up is not measured here), δ = 1, L own-scaling in the first and last module. Every layout is balanced: two KV groups and I/4 neurons per device. Contiguous is the standard layout that slices heads and neurons in storage order; random layouts are uniform balanced permutations; optimised maximises (6) with the FO score. 4

Table 1: KL to the dense model (nats per token, mean ± sample sd over seeds) and perplexity on the 16 WikiText-2 blocks after 1000 distillation steps, and the step at which the mean optimised curve first reaches the contiguous arm’s final KL. Random is eight layouts with one seed each. Layout

n

Qwen3-0.6B

optimised (FO) contiguous random

3 3 8

optimised (FO) Danube3-500M contiguous random

3 3 8

0.8

KL@500

ppl@1000

step matching contig.@1000

0.361 ± 0.009 0.303 ± 0.003 0.500 ± 0.055 0.415 ± 0.033 0.449 ± 0.044 0.374 ± 0.015

17.01 19.20 18.31

347 (2.9×) 1000 —

0.763 ± 0.056 0.667 ± 0.043 0.832 ± 0.010 0.738 ± 0.009 0.844 ± 0.023 0.737 ± 0.018

16.05 17.45 17.27

565 (1.8×) 1000 —

KL vs step: three training seeds per arm random layout (8) contiguous, mean of 3 seeds optimised (fo), mean of 3 seeds

0.7 0.6 0.5 0.4 0.3

KL@1000

200

400

600 training step

800

wikitext-2 perplexity (16 blocks)

KL to the dense model (nats/token)

Model

1000

Perplexity vs step

28 26 24 22 20 18 200

400

600 training step

800

1000

Figure 2: Qwen3-0.6B, L = 4, δ = 1: KL to the dense model (left) and perplexity (right) versus distillation step. Bold: mean of three training seeds; faint: individual seeds; grey: the eight random layouts. The Danube3-500M version is Figure 6 in the appendix.

Training. All layouts are trained with the same recipe: distillation from the frozen dense model, loss KL(teacher ∥ student) + 0.1 CE per token (Hinton et al., 2015), AdamW (Loshchilov and Hutter, 2019) with β = (0.9, 0.95), no weight decay, learning rate 5 × 10−5 with 100 warm-up steps and cosine decay to 5 × 10−6 , gradient clipping at 1, embeddings frozen, fp32 master weights with bf16 autocast. Each step is 16 sequences of 1024 tokens of FineWeb-Edu (16k tokens); the same pre-tokenised file is used for every run and the training seed only changes the block order. Short runs are 1000 steps (16M tokens), long runs 10,000 steps (164M tokens). Evaluation. Every 100 steps (250 for the long runs) we measure, on 16 blocks of 1024 tokens of the WikiText-2 test set (Merity et al., 2017), the perplexity of the DTP model and its per-token KL to the dense model in nats, which we report as the primary metric because it measures exactly what distillation minimises and does not saturate. The dense models have perplexity 12.79 (Qwen3-0.6B) and 10.70 (Danube3-500M) on these blocks. Sixteen blocks are enough for the curves and the correlations but not for differences below about 0.01 KL, and we say so where it matters. Design. Per model: three training seeds each for the optimised and contiguous layouts; eight random layouts with one seed each; seven further single-seed layouts that spread the score axis (anti-optimised, i.e. J minimised; heads only; neurons only; ADD score; and 4, 8 or 12 evenly spaced layers optimised with the rest contiguous); and one 10k-step run each for optimised and contiguous. That is 23 training runs per model, 46 in all, plus two affinity passes and sixteen layout optimisations, about 19 hours of wall clock on two RTX 5090s.

5

Results

5.1

The optimised layout trains faster and better, on every seed

Table 1 and Figure 2 give the headline. On Qwen3-0.6B the optimised layout finishes at 0.303 KL against 0.415 for contiguous and 0.374 for random, a 27% reduction against contiguous and 19%

5

Untrained (step 0) Spearman -0.16, R² 0.03

KL to dense model

add score grey: random layouts (8) 6.5 faint: individual seeds, ringed: seed mean

0.60

6.0

0.55

5.5

0.50

5.0 4.5

anti-optimisedcontiguous 8 layers 4 layers

10

15 20 fo score of the layout

optimised (fo)

contiguous

optimised (fo)

random

KL to dense model

6.25 6.00 5.75 5.50 5.25 5.00

anti-optimised

heads only 4 layers

heads only 12 layers 0.35 add score neurons only optimised (fo) 0.30

neurons only

8 10 12 fo score of the layout

optimised (fo)

contiguous

10 anti-optimised

0.70

k of 28 layers optimised, rest contiguous

After 1000 steps Spearman -0.77, R² 0.79

anti-optimised

0.75

heads only contiguous

contiguous heads only 0.70 4 layers optimised (fo) 8 layers 12 layers add score 0.65 neurons only

6 8 10 12 fo score of the layout (random ≈ 7.7, optimised 12.9) neurons only

15 20 fo score of the layout

0.80

0.90

0.75

random

heads only

After 500 steps Spearman -0.87, R² 0.82

neurons only 0.80 12 layers

6

add score

anti-optimised

optimised (fo) 0.85 add score

contiguous

4 layers contiguous 8 layers heads only 12 layers add score neurons only(fo) optimised

0.40

0.35 10 15 20 fo score of the layout (random ≈ 13.9, optimised 23.5)

0.95

anti-optimised

0.45

4 layers contiguous 8 layers

0.40

Untrained (step 0) Spearman -0.01, R² 0.00

grey: random layouts (8) 6.50 faint: individual seeds, ringed: seed mean 8 layers

After 1000 steps Spearman -0.52, R² 0.66

0.50

0.45 neurons only

heads only 12 layers

4.0

After 500 steps Spearman -0.61, R² 0.55

anti-optimised

add score

heads only

4 layers optimised (fo) 8 layers 12 layers add score neurons only

6 anti-optimised

8 10 12 fo score of the layout

k of 16 layers optimised, rest contiguous

Figure 3: KL to the dense model against the co-located FO score J of the layout (equation 6), untrained (left), after 500 (middle) and 1000 steps (right); Qwen3-0.6B top, Danube3-500M bottom. One point per layout, seeds averaged where there are several (ringed), individual seeds faint; dashed line: least-squares fit over the seventeen layouts. Untrained, the score says nothing; after training it explains most of the variance.

against random; every optimised seed is below every one of the eleven baseline runs, with a 0.05 KL clearance to the nearest random layout (Welch t-test against contiguous p = 0.027 with n = 3 per arm; against the eight random layouts p < 0.001). On Danube3-500M the margin is smaller, 0.667 against 0.738 (10%), but again all three optimised seeds are below all three contiguous seeds and below seven of the eight random layouts (Mann-Whitney against random p = 0.024; Welch against contiguous p = 0.10 at n = 3). The mean optimised curve reaches the contiguous arm’s final KL at step 347 on Qwen3-0.6B and 565 on Danube3-500M. Two remarks on the noise. Contiguous is not a privileged layout: on Danube3-500M it is indistinguishable from random (0.738 versus 0.737) and on Qwen3-0.6B it is, if anything, worse (0.415 versus 0.374), so the standard sharding is just one draw from the random distribution as far as DTP is concerned. And the larger seed spreads (contiguous on Qwen3-0.6B, optimised on Danube3-500M) are an evaluation-set property rather than a training one: in both cases the seeds’ training losses agree to three decimals, and the same seed under the 10k-step schedule reads within the tight cluster at step 1000. Sixteen blocks are noisy at the 0.03 level; the ordering holds either way. 5.2

The co-located affinity predicts the trained outcome

The objective (6) is a proxy, and the eight random layouts alone cannot test it: their scores sit within one unit of each other and their trained KLs differ by evaluation noise. The seven spread layouts (anti-optimised, layer subsets, heads only, ADD, neurons only) cover the score axis from 8.4 to 23.5 on Qwen3-0.6B (random ≈ 13.9) and 4.9 to 12.9 on Danube3-500M (random ≈ 7.7). Figure 3 plots the trained KL of all seventeen layouts against their score. After 1000 steps the Pearson correlation is −0.81 on Qwen3-0.6B (R2 = 0.66, p < 0.001) and −0.89 on Danube3-500M (R2 = 0.79); at 500 steps −0.74 and −0.91. The Spearman rank correlation is lower on Qwen3-0.6B (−0.52) because half the points are random layouts whose ranks are noise; over the nine non-random layouts it is −0.98 at both 500 and 1000 steps. The negative control matters: the anti-optimised layout, which minimises J, is the worst run in either study, 0.14 KL above random on Qwen3-0.6B and 0.09 on Danube3-500M, so the direction of the objective, not merely “structure versus none”, drives the

6

10k-step runs: KL vs step wikitext-2 perplexity

KL to dense model

0.5 0.4 0.3 0.2

s / (optimised step reaching the same KL)

training step (log)

KL gap, contiguous − optimised

Gap vs step (evals every 250 steps, 16 blocks) 17% of contig.

0.08 0.06

17% of contig.

0.04

13% of contig.

0.02 0.00

0

2000

4000 6000 training step

contiguous optimised (fo)

20 18 16 14

dense model, delta 0: 12.79

104

103

0.10

Perplexity vs step

22

contiguous optimised (fo)

16% of contig.

8000 10000

103 training step (log)

104

Compute multiplier of the optimised layout

4.0

per eval point rolling median of 5

3.5 3.0 2.5 2.0 1.5 1.0 0

2000

4000 6000 8000 10000 contiguous step s

Figure 4: Qwen3-0.6B, 10,000 distillation steps (164M tokens), one seed, evaluation every 250 steps. Top: KL and perplexity on a log step axis. Bottom left: the KL gap contiguous minus optimised with its size relative to the contiguous KL. Bottom right: the compute multiplier, the contiguous step s divided by the interpolated step at which the optimised run first reached the same KL (grey per evaluation, black rolling median of five). Danube3-500M: Figure 7.

outcome. Untrained (left panels) the correlation is zero on both models; we return to this in §5.5. 5.3

The gap does not close within 10k steps

A better initialisation could be a warm start that training washes out. Figure 4 shows the two 10k-step runs on Qwen3-0.6B. The optimised run is ahead at all 40 evaluation points. The absolute gap falls with the KL itself, from 0.088 at step 250 to 0.033 at step 10,000, but as a fraction of the contiguous KL it stays between 13% and 17% at every tabulated step (Table 2), and the last eight evaluations all read between 15.5% and 16.6%. Read as compute, the optimised run reaches each contiguous KL value 1.6 to 2.1× earlier at the tabulated steps of Table 2, and 1.5 to 2.1× earlier at every evaluation from step 1000 on, apart from a single-evaluation spike of the contiguous run at step 3750 (1.68× at step 10,000). On Danube3-500M (appendix, Figure 7) the relative gap does shrink, from 16% at step 500 to 4.6% at 10,000, but the optimised run is again ahead at all 40 points. Its multiplier is noisier: 1.6 to 2.4× at the tabulated steps, 1.3 to 3.2× per evaluation, with a rolling median that falls from 2.7× early on to 1.45× around step 7000 and ends at 1.6× (1.64× at step 10,000). Both curves are still falling at the floor learning rate, so neither number is an asymptote. The claim that both models support is that the optimised layout reaches any distillation target in about half to two thirds of the steps over the whole range we ran; whether a residual advantage survives at convergence is model-dependent and one seed per model cannot settle it. 5.4

Ablations: neurons carry the gain, heads add nothing measurable

Table 3 and Figure 5 decompose the gain. Neurons, not heads. Moving only the neurons (KV groups contiguous) captures essentially the whole effect on both models: 0.312 against 0.303 on Qwen3-0.6B and 0.620 against 0.667 on Danube3-500M, where it is in fact the best 1000-step number in the study. Moving only the KV groups (neurons contiguous) reaches a score half-way from random to optimised on Qwen3-0.6B (17.6) but trains to the random mean to three decimals (0.370), and on Danube3-500M it equals

7

Table 2: The two 10k-step runs on Qwen3-0.6B (one seed). The multiplier is the contiguous step s over the optimised step reaching the same KL. step

optimised KL

contiguous KL

gap

gap / contiguous

opt. ppl / contig. ppl

multiplier

500 1000 2000 4000 6000 8000 10000

0.392 0.333 0.281 0.237 0.207 0.186 0.174

0.470 0.397 0.325 0.283 0.247 0.218 0.207

0.079 0.064 0.043 0.046 0.040 0.032 0.033

16.7% 16.1% 13.4% 16.3% 16.3% 14.9% 16.1%

18.59 / 20.32 17.46 / 18.75 16.50 / 17.38 15.81 / 16.66 15.25 / 15.94 14.94 / 15.48 14.73 / 15.32

≥ 2.0 2.10 1.81 2.03 1.59 1.64 1.68

Table 3: Ablation layouts (single seed unless noted), ordered by Qwen3-0.6B score. Score J is the co-located FO affinity (equation 6); its random expectation is 13.75 on Qwen3-0.6B (55 edges) and 7.75 on Danube3-500M (31 edges). KL after 1000 steps. Differences below 0.01 are within evaluation noise.

Layout

What moves

anti-optimised contiguous (3 seeds) random (8 layouts) 4 layers 8 layers heads only 12 layers ADD score neurons only optimised, FO (3 seeds)

everything, J minimised nothing everything, random evenly spaced layers, incl. layer 0 KV groups; neurons contiguous everything, dot-product objective neurons; KV groups contiguous everything

Qwen3-0.6B

Danube3-500M

J

KL@1000

J

KL@1000

8.42 13.99 ≈ 13.9 15.27 16.46 17.57 18.32 21.16 22.94 23.54

0.515 0.415 0.374 0.425 0.397 0.370 0.344 0.318 0.312 0.303

4.91 7.46 ≈ 7.7 8.46 9.67 9.29 11.69 11.26 12.73 12.88

0.831 0.738 0.737 0.678 0.677 0.742 0.654 0.632 0.620 0.667

contiguous (0.742). This is consistent with the sizes of the two search spaces: the head step chooses among 2520 partitions of 8 groups, the neuron step among balanced assignments of 3072 or 4096 free variables per layer. The head step can be presented as an optional refinement; in practice one can skip it and keep the standard head sharding. First-order versus dot-product affinity. The ADD layout, rescored in FO units, reaches 87 to 90% of the optimised score and trains within 0.015 KL of it on Qwen3-0.6B and within the seed band on Danube3-500M. The first-order score is the principled one and finds a layout with a 4-point higher co-located fraction, but this study cannot separate the two objectives at the 1000-step scale. How many layers. On Danube3-500M, optimising 4 of 16 evenly spaced layers already gives 85% of the gain, 12 layers all of it; on Qwen3-0.6B the curve is close to linear in the number of layers and 4 of 28 gives nothing. The per-layer co-located fractions (appendix, Figures 9 and 10) explain the difference: on Danube3-500M the optimiser finds two layers (1 and 15) with fractions above 0.6 and the rest near random, so a subset that includes the right layer captures most of the score; on Qwen3-0.6B the optimised fraction is between 0.41 and 0.71 on every layer, with one layer at 1.00, and the score, and the gain, accumulate layer by layer. In both models the layer-subset points sit above the fit line of Figure 3: the score weights every edge equally and training does not, with the early layers counting for more than their share. The subset design confounds the count of layers with their identity; the clean test (high- versus low-fraction subsets at equal count) is three runs and is left for future work. 5.5

Zero-shot damage is uninformative

A natural shortcut would be to pick the layout with the smallest untrained perplexity. It does not work. Untrained, the optimised layout is 5th of 17 on Qwen3-0.6B (KL 4.03, perplexity 731; contiguous 4.86, 1666; random 3.9 to 6.5) and 12th of 17 on Danube3-500M (5.91, 3557; contiguous 5.02, 1506), so on one model the standard layout is the best untrained and the optimised one is below median. 8

Which component carries the gain 0.7

0.40 0.35

0.6 0.5 0.4

Layer subsets (evenly spaced, incl. layer 0) 0.44 0.42 0.40

200

400 600 training step

800

score 16.5

0.38 0.36 0.34 0.32

0.3

score 15.3 score 14.0

score 18.3

score 23.5

0.30 1000 0 4 8 12 28 number of layers whose layout is optimised (rest contiguous)

an ti

- op ti m ise co 8.4 d n ti gu 14 ous ran .0 do ≈1 m (8 he 3.9 ) ad s 17only .6 4l a 15yers .3 8l a 16yers 12 .5 la 18yers ad .3 ds c ne 21. ore uro 2 ns op 22only tim .9 ise d 23 (fo) .5

0.30

KL to dense model

KL at step 1000

0.50 0.45

contiguous (mean of 3) heads only add score neurons only optimised (fo) (mean of 3)

KL at step 1000

Every layout, with its fo score

bands: mean ± sd of the 3 optimised / 3 contiguous seeds

Figure 5: Qwen3-0.6B ablations. Left: KL at step 1000 for every layout, ordered by score (in the tick labels); shaded bands are mean ± sd of the three optimised (blue) and three contiguous (orange) seeds. Middle: KL curves of the component ablations against the two seed means. Right: KL at 1000 steps against the number of layers optimised (evenly spaced, always including layer 0).

The Spearman correlation between a layout’s untrained KL and its KL after 1000 steps is +0.19 on Qwen3-0.6B and −0.24 on Danube3-500M (neither significant), and between the score and the untrained KL it is −0.16 and −0.01 (Figure 3, left panels; appendix Figure 11). By step 100, the end of warm-up, the optimised run is already ahead on both models and the order never flips again. The zero-shot damage of an untrained DTP model is dominated by a few large, easily repaired disruptions; what the layout controls is how much of the network’s fine structure the devices can reconstruct once those are repaired. 5.6

Where the affinity concentrates

The FO score is not uniformly informative across the network. On Qwen3-0.6B the optimiser drives the co-located fraction of layer 2’s up edge to 1.00 (the anti-optimised layout drives it to 0.00), while ADD reaches only 0.46 there. Probing the dense model shows why: six FFN neurons of layer 2 write an activation of roughly 7000 into a single residual dimension on delimiter tokens, driven almost entirely by two query heads of one KV group. This is the massive-activation circuit described by Sun et al. (2024), which originates in the early-layer FFN on delimiter tokens and feeds the attention sinks of later layers (Sun et al., 2026; Xiao et al., 2024; Yu et al., 2024). The first-order score sees the circuit because it measures the damage in residual-stream units through the SwiGLU operating point; the pre-activation dot product does not, because the neurons’ pre-activations are not unusual, only their gain is. More generally, the affinity matrices are far from uniform: averaged over neurons, the single most influential KV group carries 27 to 64% of a neuron’s up-edge affinity (Qwen3-0.6B, per layer; 27 to 48% on Danube3-500M) against 12.5% if the eight groups contributed equally, which is the structure that Knittel and Pfister (2026) report as sparse inter-layer dependencies of FFN neurons on attention outputs, and that Neo et al. (2024) describe for next-token neurons driven by specific heads.

6

Related work

Architectures that communicate less. Parallel attention and FFN blocks (Chowdhery et al., 2023; Wang and Komatsuzaki, 2021) halve the all-reduces per layer. Ladder Residual (Zhang et al., 2025) routes each module’s input from the previous module’s residual so the all-reduce overlaps with compute, and converts Llama-3.1-8B with 3B tokens of retraining. Kraken (Prabhakar et al., 2024) and Parallel Track Transformers (Wang et al., 2026) run several thinner Transformers in parallel and fuse them periodically, reducing syncs by up to 16×; CAAT-Net (Lamprecht et al., 2025) allreduces only a subset of channels. DTP (Kog Team, 2026) keeps the dense Transformer’s width and per-module structure and delays the reduction instead. All of these leave each device with partial information for part of the forward pass, and none chooses which heads and neurons a device holds; our contribution is that choice, and it applies in principle to any of them. Sync-Point Drop. The closest prior work is SPD (Kim et al., 2025), which deletes the attention all-reduce in blocks that tolerate it and, for the most sensitive blocks, re-initialises the sharding before 9

block-to-block distillation. SPD’s initialisation scatters heads across devices by maximising the distance between the attention patterns of heads that share a device, so that every device holds a functionally diverse set, and then matches each head group with one of the existing MLP partitions by output norm. Ours differs in objective and granularity: we co-locate heads with the neurons that depend on them, we place individual neurons rather than whole partitions, and our ablations show the neuron assignment is what carries the gain. SPD reports a 3% accuracy recovery from its head grouping on LLaMA2-7B; a direct comparison under DTP is future work. Systems approaches. Kernel-level overlap (Chang et al., 2024; Wang et al., 2024), faster all-reduce primitives (Taneja et al., 2026) and compressed activations (Hansen-Palmus et al., 2024; Li et al., 2024) reduce the cost of the collective without changing the model. They compose with architectural changes such as DTP and with the layout choice studied here. Permutation symmetry and hardware-aware permutations. That hidden units can be permuted without changing the function underlies weight matching for model merging (Ainsworth et al., 2023; Entezari et al., 2022); Tran et al. (2026) characterise the corresponding symmetries of attention heads under rotary embeddings, which is the head-level symmetry we use. Pool and Yu (2021) permute channels so that a network fits the N:M sparsity pattern of the hardware without accuracy loss; our work is the same move for a communication constraint. Neuron clustering and modularity. MoEfication (Zhang et al., 2022), emergent modularity (Zhang et al., 2023) and LLaMA-MoE (Zhu et al., 2024) partition FFN neurons into experts by co-activation or weight similarity to enable conditional computation. We partition neurons by their dependence on specific heads, for co-location rather than routing, and every neuron still runs on every token. The key-value-memory view of the FFN (Geva et al., 2021) and the sparse attention-to-neuron dependencies of Knittel and Pfister (2026) are the structure that makes a good partition exist. Converting dense checkpoints by distillation. Uptraining a pretrained model into a cheaper architecture with a short distillation is standard practice (Ainslie et al., 2023; Kim et al., 2025; Muralidharan et al., 2024; Zhang et al., 2025). Our result is about the starting point of that distillation: a better permutation of the same weights roughly halves the number of steps to a target.

7

Limitations and future work

Everything here is L = 4 and δ = 1 on two sub-billion-parameter models, and the devices are simulated on one GPU, so we make no claim about wall-clock speed. At larger δ a module misses the other devices’ partials of the last δ modules and the affinity graph gains edges; the objective and the optimiser extend directly, but we have not run it. The seed studies have n = 3 per arm and the long runs one seed, evaluated on 16 blocks of WikiText-2; the direction of every comparison is consistent, but differences below 0.01 KL should be read as ties, and the final numbers deserve the full test set and downstream tasks. The layer-subset ablation confounds how many layers move with which ones. Finally, the affinity is measured on the dense model and used once; re-estimating it on the partially trained DTP model, or making the layout part of training, are natural extensions. A shared-expert variant that replicates 10% of each layer’s neurons on every device gave no further gain in a preliminary run (appendix).

8

Conclusion

Delayed Tensor Parallelism turns the sharding of a Transformer from a systems detail into a modelling choice: a device can only see the previous module through the heads and neurons it holds. Because heads and neurons are permutation-symmetric, the choice is free. A first-order affinity from one calibration pass and a two-minute assignment optimiser produce a layout that reaches any distillation target in about half the steps of the standard layout, on both models we tried, and the objective it maximises predicts the trained outcome across seventeen layouts. The gain comes from placing FFN neurons next to the heads they depend on. We expect the same principle to apply to every architecture that trades synchronisation for partial information.

10

Acknowledgements We thank the Kog team for the DTP design and for discussions. A large language model was used to draft the text and figures of this paper from the author’s notes and experiment logs; the author checked every statement, number and reference.

References Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. GQA: Training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023. Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. Git re-basin: Merging models modulo permutation symmetries. In International Conference on Learning Representations (ICLR), 2023. Li-Wen Chang, Wenlei Bao, Qi Hou, Chengquan Jiang, Ningxin Zheng, Yinmin Zhong, Xuanrun Zhang, Zuquan Song, Ziheng Jiang, Haibin Lin, Xin Jin, and Xin Liu. FLUX: Fast software-based communication overlap on GPUs through kernel fusion, 2024. arXiv:2406.06858. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023. David F. Crouse. On implementing 2D rectangular assignment algorithms. IEEE Transactions on Aerospace and Electronic Systems, 52(4):1679–1696, 2016. Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, https: //transformer-circuits.pub/2021/framework/index.html, 2021. Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur. The role of permutation invariance in linear mode connectivity of neural networks. In International Conference on Learning Representations (ICLR), 2022. Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021. Jan Hansen-Palmus, Michael Truong Le, Oliver Hausdörfer, and Alok Verma. Communication compression for tensor parallel LLM inference, 2024. arXiv:2411.09510. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network, 2015. arXiv:1503.02531. Han-Byul Kim, Duc Hoang, Arnav Kundu, Mohammad Samragh, and Minsik Cho. SPD: Syncpoint drop for efficient tensor parallelism of large language models. In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025. arXiv:2502.20727. Johannes Knittel and Hanspeter Pfister. Sparse inter-layer dependencies of transformer FFN neurons, 2026. arXiv:2607.11990. Kog Team. Delayed tensor parallelism for faster transformer inference, May 2026. Kog Labs blog post. https://blog.kog.ai/delayed-tensor-parallelism-forfaster-transformer-inference/.

11

Harold W. Kuhn. The Hungarian method for the assignment problem. Naval Research Logistics Quarterly, 2:83–97, 1955. Itay Lamprecht, Asaf Karnieli, Yair Hanani, Niv Giladi, and Daniel Soudry. Tensor-parallelism with partially synchronized activations. In Advances in Neural Information Processing Systems 38 (NeurIPS), 2025. arXiv:2506.19645. Qingyuan Li, Bo Zhang, Liang Ye, Yifan Zhang, Wei Wu, Yerui Sun, Lin Ma, and Yuchen Xie. Flash communication: Reducing tensor parallelization bottleneck for fast large language model inference, 2024. arXiv:2412.04964. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019. Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations (ICLR), 2017. Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, and Pavlo Molchanov. Compact language models via pruning and knowledge distillation. In Advances in Neural Information Processing Systems 37 (NeurIPS), 2024. Clement Neo, Shay B. Cohen, and Fazl Barez. Interpreting context look-ups in transformers: Investigating attention-MLP interactions. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. Guilherme Penedo, Hynek Kydlíček, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro von Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. In Advances in Neural Information Processing Systems 37 (NeurIPS), Datasets and Benchmarks Track, 2024. Pascal Pfeiffer, Philipp Singer, Yauhen Babakhin, Gabor Fodor, Nischay Dhankhar, and Sri Satish Ambati. H2O-Danube3 technical report, 2024. arXiv:2407.09276. Jeff Pool and Chong Yu. Channel permutations for N:M sparsity. In Advances in Neural Information Processing Systems 34 (NeurIPS), 2021. Rohan Baskar Prabhakar, Hengrui Zhang, and David Wentzlaff. Kraken: Inherently parallel transformers for efficient multi-device inference. In Advances in Neural Information Processing Systems 37 (NeurIPS), 2024. arXiv:2408.07802. Qwen Team. Qwen3 technical report, 2025. arXiv:2505.09388. Noam Shazeer. GLU variants improve transformer, 2020. arXiv:2002.05202. Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism, 2019. arXiv:1909.08053. Mingjie Sun, Xinlei Chen, J. Zico Kolter, and Zhuang Liu. Massive activations in large language models. In First Conference on Language Modeling (COLM), 2024. arXiv:2402.17762. Shangwen Sun, Alfredo Canziani, Yann LeCun, and Jiachen Zhu. The spike, the sparse and the sink: Anatomy of massive activations and attention sinks, 2026. arXiv:2603.05498. Hritvik Taneja, Anish Saxena, Abhishek Revinipati, Jae Hyung Ju, Neal C. Crago, and Moinuddin Qureshi. SiFAR: Synchronization-free all-reduce for low-latency LLM inference, 2026. arXiv:2607.08973.

12

Viet-Hoang Tran, Vinh Khanh Bui, Van-Hoan Trinh, Tan Lai Ngoc, and Tan M. Nguyen. Functional equivalence in attention: A comprehensive study with applications to linear mode connectivity, 2026. arXiv:2606.17830. Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 billion parameter autoregressive language model. https://github.com/kingoflolz/mesh-transformer-jax, 2021. Chong Wang, Nan Du, Tom Gunter, Tao Lei, Kulin Seth, Senyu Tong, Jianyu Wang, Guoli Yin, Xiyou Zhou, Kelvin Zou, and Ruoming Pang. Parallel track transformers: Enabling fast GPU inference with reduced synchronization, 2026. arXiv:2602.07306. Guanhua Wang, Chengming Zhang, Zheyu Shen, Ang Li, and Olatunji Ruwase. Domino: Eliminating communication in LLM training via generic tensor slicing and overlapping, 2024. arXiv:2409.15241. Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. In International Conference on Learning Representations (ICLR), 2024. Mengxia Yu, De Wang, Colorado Reed, and Alvin Wan. The super weight in large language models, 2024. arXiv:2411.07191. Muru Zhang, Mayank Mishra, Zhongzhu Zhou, William Brandon, Jue Wang, Yoon Kim, Jonathan Ragan-Kelley, Shuaiwen Leon Song, Ben Athiwaratkun, and Tri Dao. Ladder-residual: Parallelismaware architecture for accelerating large model inference with communication overlapping. In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025. arXiv:2501.06589. Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. MoEfication: Transformer feed-forward layers are mixtures of experts. In Findings of the Association for Computational Linguistics: ACL 2022, 2022. Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Chaojun Xiao, Xiaozhi Wang, Xu Han, Zhiyuan Liu, Ruobing Xie, Maosong Sun, and Jie Zhou. Emergent modularity in pre-trained transformers. In Findings of the Association for Computational Linguistics: ACL 2023, 2023. Tong Zhu, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, and Yu Cheng. LLaMAMoE: Building mixture-of-experts from LLaMA with continual pre-training. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024.

A

Reproducibility

Code, layouts and logs are in the accompanying repository. Per model the whole study is one generated shell chain: pre-tokenise FineWeb-Edu, run the affinity pass, build the eight optimised or ablation layouts and the eight random layouts, then the 23 training runs; every run is skipped if its log is complete, so the chain is resumable. The figures and tables are produced from the logs by one script. Hyper-parameters not stated in §4: micro-batch 4 with 4 accumulation steps; gradient checkpointing on; calibration and training tokens from the same pre-tokenised file with the first 64 blocks reserved for calibration; random layouts drawn with fixed seeds 1000 to 1007; the optimiser’s neuron assignment uses 300 dual-ascent iterations with step decay 0.98 before the fix-up and swap phases.

B

Earlier runs and the shared-expert variant

Before the controlled study we ran a 2000-step version of the same recipe on Qwen3-0.6B with one optimised and one random layout (perplexity 16.42 and 17.82, KL 0.265 and 0.346 at 2000 steps; the optimised run reached the random run’s final perplexity at step 814). A variant in which 10% of each layer’s neurons are replicated on every device, added locally and never broadcast, finished at 16.26 / 13

0.260 under the same protocol, within noise of the plain optimised layout, so we did not pursue it in the controlled study.

Additional figures: Danube3-500M

KL vs step: three training seeds per arm

Perplexity vs step

random layout (8) contiguous, mean of 3 seeds optimised (fo), mean of 3 seeds

1.4

wikitext-2 perplexity (16 blocks)

KL to the dense model (nats/token)

C

1.2 1.0 0.8 0.6

200

400

600 training step

800

1000

35 30 25 20 15

200

400

600 training step

800

1000

Figure 6: Danube3-500M seed curves, as Figure 2.

10k-step runs: KL vs step wikitext-2 perplexity

KL to dense model

0.9 0.8 0.7 0.6 0.5

KL gap, contiguous − optimised

Gap vs step (evals every 250 steps, 16 blocks) 16% of contig.

9% of contig.

s / (optimised step reaching the same KL)

training step (log)

10% of contig.

5% of contig.

0

2000

4000 6000 training step

contiguous optimised (fo)

20 18 16 14 12

dense model, delta 0: 10.70

104

103

0.175 0.150 0.125 0.100 0.075 0.050 0.025 0.000

Perplexity vs step

22

contiguous optimised (fo)

8000 10000

103 training step (log)

104

Compute multiplier of the optimised layout per eval point rolling median of 5

3.0 2.5 2.0 1.5 1.0 0

2000

4000 6000 8000 10000 contiguous step s

Figure 7: Danube3-500M 10k-step runs, as Figure 4. The relative gap falls from 16% to 4.6% but the optimised run is ahead at all 40 evaluations; the multiplier is 1.6 to 2.4× at the tabulated steps and its rolling median of five stays between 1.45 and 2.7×.

14

Which component carries the gain contiguous (mean of 3) heads only add score neurons only optimised (fo) (mean of 3)

1.2 KL to dense model

0.80 KL at step 1000

1.3

0.75 0.70 0.65

1.1 1.0 0.9 0.8

0.72 0.70 0.68 0.66

score 8.5 score 9.7

score 12.9 score 11.7

200

400 600 training step

800

0.62 1000 0 4 8 12 16 number of layers whose layout is optimised (rest contiguous)

an ti

- op ti m ise co 4.9 d n ti gu ou ran 7.5 s do m ≈7 (8) he .7 ad so 9.3nly 4l ay e 8.5 rs 8l ay e 9 rs 12 .7 lay e 11 rs ad .7 ds c ne 11. ore uro 3 ns op 12only tim .7 ise d 12 (fo) .9

0.6

score 7.5

0.64

0.7

0.60

Layer subsets (evenly spaced, incl. layer 0) 0.74 KL at step 1000

Every layout, with its fo score

bands: mean ± sd of the 3 optimised / 3 contiguous seeds

0.85

Figure 8: Danube3-500M ablations, as Figure 5. Neurons-only is the best 1000-step layout; heads-only equals contiguous; four evenly spaced layers give 85% of the gain.

D

Per-layer co-located fractions Co-located fraction of the heads → same-layer-neurons edge, per layer (random = 0.25)

optimised (fo) (23.5)

0.54 0.60 1.00 0.60 0.54 0.52 0.57 0.71 0.45 0.50 0.44 0.53 0.41 0.49 0.48 0.49 0.47 0.53 0.52 0.48 0.63 0.55 0.61 0.61 0.60 0.53 0.57 0.44

neurons only (22.9) 0.51 0.60 1.00 0.60 0.54 0.52 0.57 0.69 0.47 0.47 0.45 0.48 0.41 0.49 0.48 0.49 0.47 0.53 0.52 0.48 0.63 0.56 0.58 0.61 0.60 0.53 0.57 0.40

1.0 0.8

add score (21.2) 0.43 0.38 0.46 0.43 0.37 0.34 0.41 0.34 0.36 0.40 0.37 0.39 0.34 0.40 0.39 0.44 0.42 0.44 0.44 0.38 0.44 0.34 0.37 0.31 0.34 0.35 0.34 0.32 12 layers (18.3) 0.54 0.26 1.00 0.28 0.26 0.52 0.23 0.71 0.26 0.50 0.24 0.33 0.41 0.20 0.48 0.27 0.47 0.28 0.21 0.48 0.30 0.56 0.30 0.61 0.38 0.33 0.57 0.20

0.6

heads only (17.6) 0.36 0.29 0.74 0.36 0.36 0.33 0.42 0.60 0.30 0.31 0.34 0.42 0.31 0.37 0.39 0.29 0.33 0.31 0.33 0.30 0.36 0.34 0.33 0.40 0.39 0.34 0.43 0.38

0.4

8 layers (16.5) 0.54 0.26 0.26 0.28 0.54 0.25 0.23 0.71 0.26 0.23 0.45 0.33 0.24 0.20 0.48 0.27 0.29 0.28 0.54 0.23 0.30 0.56 0.30 0.19 0.60 0.33 0.21 0.20 4 layers (15.3) 0.54 0.26 0.26 0.28 0.26 0.25 0.23 0.71 0.26 0.23 0.24 0.33 0.24 0.20 0.48 0.27 0.29 0.28 0.21 0.23 0.30 0.56 0.30 0.19 0.38 0.33 0.21 0.20 anti-optimised (8.4) 0.12 0.09 0.00 0.09 0.10 0.12 0.09 0.07 0.12 0.14 0.14 0.12 0.15 0.13 0.12 0.13 0.12 0.10 0.12 0.12 0.07 0.09 0.08 0.09 0.09 0.10 0.09 0.11

0.2 0.0

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer

Figure 9: Qwen3-0.6B: co-located fraction of the up edge (heads to same-layer neurons) per layer for each optimised layout, random = 0.25. Layer 2, which hosts the massive-activation circuit of §5.6, is fully separable: the optimiser co-locates 100% of its affinity, the anti-optimised layout 0%, and the head sweep alone already reaches 0.74 with contiguous neurons. Eight layers reach 0.6 or more and none is below 0.41.

15

Co-located fraction of the heads → same-layer-neurons edge, per layer (random = 0.25) optimised (fo) (12.9)

0.47

0.92

0.52

0.54

0.51

0.42

0.52

0.45

0.45

0.42

0.46

0.49

0.50

0.53

0.52

0.87

neurons only (12.7)

0.47

0.78

0.52

0.54

0.51

0.42

0.52

0.45

0.45

0.42

0.46

0.49

0.50

0.53

0.52

0.87

12 layers (11.7)

0.47

0.92

0.25

0.54

0.51

0.42

0.25

0.45

0.45

0.42

0.27

0.49

0.50

0.53

0.23

0.90

add score (11.3)

0.59

0.39

0.41

0.38

0.40

0.35

0.44

0.37

0.39

0.38

0.36

0.41

0.38

0.38

0.53

0.61

8 layers (9.7)

0.47

0.18

0.52

0.23

0.51

0.24

0.52

0.23

0.45

0.26

0.46

0.24

0.50

0.25

0.52

0.07

heads only (9.3)

0.28

0.92

0.27

0.29

0.26

0.27

0.27

0.27

0.26

0.26

0.27

0.27

0.27

0.27

0.34

0.73

4 layers (8.5)

0.47

0.18

0.25

0.23

0.51

0.24

0.25

0.23

0.45

0.26

0.27

0.24

0.50

0.25

0.23

0.07

anti-optimised (4.9)

0.13

0.01

0.10

0.11

0.12

0.15

0.11

0.14

0.14

0.15

0.13

0.10

0.11

0.10

0.12

0.01

0

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

layer

1.0 0.8 0.6 0.4 0.2 0.0

Figure 10: Danube3-500M: as Figure 9. Layers 1 and 15 are the only ones above 0.6; layer 1 is the only layer where the head sweep adds to the neuron assignment (0.78 with contiguous heads, 0.92 with optimised heads).

E

Untrained versus trained

The ordering flips within the first 100 steps

Untrained vs trained KL, all 1000-step runs (Spearman +0.19) anti-optimised contiguous (3 seeds) 4 layers 8 layers

0.45 0.40 0.35 0.30

KL (log)

KL at step 1000

0.50

heads only 12 layers

4 5 6 KL at step 0 (untrained DTP model)

contiguous (3 seeds)

0.70

4 layers 12 layers

heads only 8 layers

KL (log)

KL at step 1000

200

400 600 training step

800

1000

contiguous s1 optimised s1 random5 (best untrained) random7 (worst untrained)

6 × 100

0.80

0.65

0

The ordering flips within the first 100 steps

Untrained vs trained KL, all 1000-step runs (Spearman -0.24) anti-optimised 0.75

contiguous s1 optimised s1 random5 (best untrained) random7 (worst untrained)

100 6 × 10−1 4 × 10−1 3 × 10−1

add score

neurons only optimised (3 seeds)

6 × 100 4 × 100 3 × 100 2 × 100

4 × 100 3 × 100 2 × 100 100

add score neurons only optimised (3 seeds)

5.0 5.5 6.0 6.5 KL at step 0 (untrained DTP model)

0

200

400 600 training step

800

1000

Figure 11: Left: KL at step 0 against KL at step 1000 for every 1000-step run (Qwen3-0.6B top, Danube3-500M bottom). Right: full KL trajectories on a log scale for the contiguous and optimised seed-1 runs and the best and worst random layouts untrained. On Danube3-500M the contiguous layout is the best untrained layout of the seventeen and the optimised one is twelfth; by step 100 the order has flipped and it never flips back.

16

Record · ID 919373 · SHA-256 95703552c39e6c3d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.