ConceptioArchivearXiv CS
arXiv CSopen access

A Mechanistic Analysis of Looped Reasoning Language Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

A Mechanistic Analysis of Looped Reasoning Language Models

Hugh Blayney 1 Álvaro Arroyo 1 Johan Obando-Ceron 2 3 Pablo Samuel Castro 2 3 Aaron Courville 2 3 Michael Bronstein 1 4 Xiaowen Dong 1 Recurrences 0-16

Recurrences 16-32

10

10

5

5

0

0

-5

-5

-10

Late Recurrence

Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM’s layers in the latent dimension, resulting in looped reasoning language models. Despite promising results, few works have investigated how their internal dynamics differ from those of standard feedforward models. In this paper, we conduct a mechanistic analysis of the latent states in looped language models, focusing in particular on how the stages of inference observed in feedforward models compare to those observed in looped ones. To this end, we analyze cyclic recurrence and show that for many of the studied models each layer in the cycle converges to a distinct fixed point; consequently, the recurrent block follows a consistent cyclic trajectory in the latent space. We provide evidence that as these fixed points are reached, attention-head behavior stabilizes, leading to constant behavior across recurrences. Empirically, we discover that recurrent blocks learn stages of inference that closely mirror those of feedforward models, repeating these stages in depth with each iteration. We study how recurrent block size, input injection, and normalization influence the emergence and stability of these cyclic fixed points. We believe these findings help translate mechanistic insights into practical guidance for architectural design.

PC 2

arXiv:2604.11791v1 [cs.LG] 13 Apr 2026

Abstract

-10 -10

0 PC 1

10

-10

0

10

Early

PC 1

Figure 1. Latent states after each block in a recurrent model frequently tend towards separate fixed points, meaning that the application of a recurrent block tends towards a consistent trajectory in latent space.

through chain-of-thought (CoT) prompting (Wei et al., 2022) or reinforcement-learning-based fine-tuning, first popularized in the DeepSeek-R1 architecture (Guo et al., 2025). More recently, research has explored building reasoning capabilities directly into the model architecture via recurrent looping (Geiping et al., 2025; Wang et al., 2025; JolicoeurMartineau, 2025), where additional test-time compute is spent by taking more recurrent steps, echoing early designs in this direction (Graves, 2016). Despite growing empirical success, the mechanisms underlying these models remain poorly understood, as well as their benefits and limitations when compared to feedforward computation. In this paper, we compare how feedforward and looped LLMs organize computation across (effective) depth through the lens of stages of inference (Lad et al., 2024; Queipo-de Llano et al., 2025), a perspective suggesting that LLM inference can be decomposed into several distinct computational stages. Building on prior observations that repeated application of a shared recurrent block can approach a fixed point or steady state (Yang et al., 2023; Geiping et al., 2025), we show that such behavior necessarily implies one of two possibilities: either the contribution of the component Transformer blocks vanishes asymptotically, or their sequential application traces out a constant cyclic trajectory in latent space. We further demonstrate empirically that the latter behavior arises in practice when certain architectural conditions are met, and that this appears to be emergent behavior from the Transformer architecture itself, appearing in both trained recurrent models and untrained, randomly initialized models.

1. Introduction The vast majority of current LLMs are based on the Transformer architecture (Vaswani et al., 2017), which comprises a sequence of blocks traversed in a feedforward manner to predict the next token. As the capability of these models increased, attention turned to eliciting reasoning capabilities in LLMs by increasing test-time computation, commonly 1 University

of Oxford 2 Mila – Quebec AI Institute de Montréal 4 AITHYRA. Correspondence to: Hugh Blayney <[email protected]>, Álvaro Arroyo <[email protected]>. 3 Université

1

A Mechanistic Analysis of Looped Language Models

In the case of Geiping et al. (2025); McLeish et al. (2025) this stacked block may also take an additional input Z ∈ R𝑇 ×𝐷 which is typically initialized from a Normal distribution, as well as the original input to the recurrent section: this is known as input injection (Bai et al., 2019; Anil et al., 2022), and the two inputs are projected into common feature space R𝐷 before the block is applied. In the case of input injection, a 𝑘-stacked block therefore becomes

Our analysis yields two key insights: 1. Sec. 4 establishes that many looped language models tend toward cyclic fixed-point behavior, providing theoretical and empirical proof that this implies convergence to fixed attention patterns. Sec. 4.2 explores how specific architectural choices influence this behavior. 2. Sec. 5 demonstrates that these stable attention patterns mirror the “mixing” stages of inference learned in feedforward models. We then provide evidence in Sec. 5.1 that models naturally self-organize into these stages during training, and show in Sec. 5.2 that stable fixed-point models successfully maintain these inference stages, whereas unstable models deviate.

S 𝑘 (X, Z) = B 𝑘 (B 𝑘−1 (. . . B1 ( [X, Z]W𝐼 ) . . . )),

where [·, ·] denotes concatenation in the channel dimension and W𝐼 ∈ R2𝐷×𝐷 is a learned projection matrix. This allows us to define a (𝑘 ⊗ 𝑙)-Recurrent block as a 𝑘-stacked block repeated 𝑙 times: ×𝑙

z }| { 𝑅𝑙,𝑘 (X) = S 𝑘 (S 𝑘 (. . . S 𝑘 ( X) . . . )),

2. Preliminaries

(6)

which with input-injection becomes

2.1. Looped Transformers

×𝑙

z }| { 𝑅𝑙,𝑘 (X, Z) = S 𝑘 (X, S 𝑘 (. . . X, S 𝑘 (X, Z)) . . . )) .

In this section we introduce Looped Transformers and define the notation that we will use throughout the paper. We represent the input sequence to our transformer of length 𝑇 and dimension 𝐷 as X ∈ R𝑇 ×𝐷 . Following the notation of Yudin et al. (2025) we define Attn, the dot product selfattention mechanism, as a map 𝑓 : R𝑇 ×𝐷 → R𝑇 ×𝐷 :   XW𝑄 W𝐾⊤ X ⊤ Attn(X) = softmax XW𝑉 , (1) √ 𝑑 = softmax( 𝐴(X))XW𝑉 , (2)

(7)

Note that the input X is only “injected” at the start of each stack of blocks – once per recurrence. Additionally, a complete looped Transformer may have multiple feedforward layers before the Recurrent block, and multiple feedforward layers after the Recurrent block: following the convention of Geiping et al. (2025), we refer to these as prelude and coda layers respectively, and these are simply separate stacked blocks with non-tied layer weights. Where prelude and coda are used, we will frequently refer to this as a sandwich block structure.

where W𝑄 , W𝐾 , W𝑉 ∈ R𝐷×𝑑 are projection matrices and 𝐴 is defined for convenience.

We combine and adapt the nomenclature of Geiping et al. (2025); Saunshi et al. (2025) and refer to a looped Transformer with 𝑝 prelude layers, 𝑘 recurrent layers and 𝑐 coda layers with the tuple ( 𝑝, 𝑘, 𝑐). Where input injection is used, we add an 𝐼 subscript ( 𝑝, 𝑘, 𝑐) 𝐼 , and when referring to a recurrent layer looped a specific number of times 𝑙, we denote this as ( 𝑝, 𝑘 ⊗ 𝑙, 𝑐).

A transformer block typically comprises an attention mechanism and a position-wise MLP as follows: X̂ = 𝑛2 (X + Attn(𝑛1 (X)) ,   X ′ = 𝑛4 X̂ + MLP(𝑛3 ( X̂)) ,

(5)

(3) (4)

where 𝑛1 , 𝑛2 , 𝑛3 , 𝑛4 are each optional norms – here we are borrowing from the notation of Geiping et al. (2025). We denote the action of a Transformer block B : R𝑇 ×𝐷 → R𝑇 ×𝐷 as X ′ = B(X), and refer to the intermediate hiddenstate matrices X between blocks as the residual stream.

In summary, a ( 𝑝, 𝑘 ⊗ 𝑙, 𝑐) looped Transformer is defined as X0 ← S 𝑝 (X) X𝑖 ← S′𝑘 (X𝑖−1 )

𝑖 ∈ {1, . . . , 𝑙}

X ← S′′ 𝑐 (X𝑙 ),

Looped Transformers are Transformers that utilize “recurrence in depth” – that is, they reapply layers to repeatedly act on the latent states. Recent research has identified that an effective way (Bae et al., 2025) to achieve this is via cyclic recurrence: a fixed sequence of layers is repeated in a “cyclic” pattern. This is the approach that we focus on in this work, and we introduce it in more detail below. For convenience, we define a 𝑘-stacked block as a composition of Transformer blocks, S 𝑘 (X) = B 𝑘 (B 𝑘−1 (. . . B1 (X) . . . )).

where ′ and ′′ indicates that these are different stacks between which weights are not shared. A ( 𝑝, 𝑘 ⊗ 𝑙, 𝑐) 𝐼 looped Transformer (with input injection) is defined as X ← S 𝑝 (X) Z𝑖 ← S′𝑘 (X, Z𝑖−1 ) X ← S′′ 𝑐 (Z𝑙 ), 2

𝑖 ∈ {1, . . . , 𝑙}

A Mechanistic Analysis of Looped Language Models

3. Related Work

where Z0 is initialized such that each column is sampled from N (0, 𝜎 2 I𝐷 ).

Looped and Recurrent Transformers Reusing the same Transformer block for multiple iterations is an idea that has been explored in the literature. This began with the introduction of Universal Transformers (Dehghani et al., 2018), which have also resulted in sparsified and conditionalcomputation extensions (Tan et al., 2023; Csordás et al., 2024). Other more recent recurrent style architectures with a higher focus on reasoning-style tasks have been HRM (Wang et al., 2025) and TRM (Jolicoeur-Martineau, 2025). Within language modeling, we highlight Huginn0125 (Geiping et al., 2025), Ouro (Zhu et al., 2025), and Mixture-of-Recursions (Bae et al., 2025) as models that have been pretrained from random initialization, as well as recent work by (McLeish et al., 2025; Koishekenov et al., 2025) that retrofit recurrence into pretrained LLMs.

We will describe our results via grouping: • No grouping: The value is visualized as it evolves through sequential layers of the model, irrespective of whether these layers are repeated. • By recurrence: Separate lines are visualized for each complete pass through the recurrent block. The 𝑥axis is typically percentage scaled to represent relative depth within that block (including prelude/coda), allowing us to overlay and compare successive passes. Always presented with a green-yellow colorbar, with later recurrences colored more yellow. • By layer: Separate lines are visualized for each unique layer, showing how the value evolves across recurrences. Always presented with a blue-green colorbar, with later layers colored more green.

In terms of mechanistic studies, we highlight Pappone et al. (2025), who analyze “two-scale” latent dynamics in recurrent Transformers. However, the setting of their analysis is different from ours: they analyse a model in which each recurrent block comprises either 1 or 2 layers, and the model comprises multiple separate recurrent blocks. In this way, the two scales they refer to correspond to the outputs of each iteration of a given recurrent block, and outputs of recurrent blocks when transitioning between different recurrent blocks. We instead study a single looped block with deeper cycles (4+ layers), closer to common looped architectures, and we analyze the internals of these cyclic blocks by examining the latent states of each separate layer. We also highlight work related to looped model expressivity (Xu & Sato, 2024; Saunshi et al., 2025), as well as work on neural network stability and fixed-point dynamics (Bai et al., 2019; Anil et al., 2022; Ke et al., 2024; Yudin et al., 2025). The only work – of which we are aware – that analyses the internal states of the cyclic recurrent blocks is Lu et al. (2025), who demonstrate cyclic behavior in logit lens prediction throughout recurrent blocks.

2.2. Stages of Inference The behavior of layers in feedforward Transformers appears to change sharply with depth: Lad et al. (2024) originate the term “stages of inference” and demonstrate how several different layer mechanisms emerge at different depths. Queipo-de Llano et al. (2025) further develop this viewpoint, focusing on behaviors that can be characterized by the mixing (or lack thereof) induced by the attention heads. We focus on this latter perspective. Mixing in this context refers to the extent to which the attention mechanism incorporates information from previous tokens at each layer. Throughout the main text of this paper we quantify our study of mixing behavior through the ColSum Concentration metric introduced in Queipo-de Llano et al. (2025); we introduce and discuss additional metrics in App. E. Í We first define the column sum 𝑐 𝑗 = 𝑖 𝐴𝑖 𝑗 to capture how much attention mass is received by token 𝑗. We normalize this as Í 𝑐ˆ 𝑗 = 𝑐 𝑗 /𝑇 to obtain a probability distribution, noting that 𝑖, 𝑗 𝐴𝑖 𝑗 = 𝑇 since 𝐴 is row-stochastic. From this distribution, we define the ColSum Concentration via its normalÍ ized entropy as 𝐶 = 1−𝐻col ∈ [0, 1] = 1+ log1 𝑇 𝑗 𝑐 𝑗 log 𝑐 𝑗 .

Stages of Inference and Attention Dynamics The idea that LLMs organize their feedforward computation into several distinct stages of inference was first proposed by Lad et al. (2024). Building on this, Queipo-de Llano et al. (2025) explain the emergence of these stages through the behavior of attention heads, driven by massive activations (Sun et al., 2024). We also highlight a complementary line of work that analyzes LLM learning dynamics via attention patterns through the lens of mixing (Barbero et al., 2024; Arroyo et al., 2026; Veličković et al., 2024; Barbero et al., 2025), motivated by information propagation challenges originally studied in Graph Neural Networks (GNNs) (Cai & Wang, 2020; Alon & Yahav, 2020; Arroyo et al., 2025; Hariri et al., 2025; Blayney et al., 2025).

Large values of 𝐶 indicate a high concentration of attention mass: few columns receive most of the mass. We note therefore that this metric captures the well-studied attention sink (Xiao et al., 2023; Barbero et al., 2025) behavior, but generalizes to capture concentration over any token position. This property is useful for our investigation as not all models studied herein exhibit sinks on the first token; in particular OLMo-2 frequently concentrates attention mass on punctuation, echoing a result in Sandoval-Segura et al. (2025). 3

A Mechanistic Analysis of Looped Language Models

bounded2 – this implies that the attention patterns will converge, as shown in Proposition 4.2.

Test-time Computation Test-time computation broadly refers to giving a model the ability to expend additional computational cycles at inference in proportion to the difficulty of the input, rather than committing to a fixed compute budget for all examples. In this paper, we focus specifically on recurrence as a mechanism for scaling computation at test time, as opposed to alternative strategies such as early-exit architectures (Schuster et al., 2022) or continuous thought machines (Darlow et al., 2025). Classic approaches to adaptive test-time compute include Adaptive Computation Time (Graves, 2016) and subsequent probabilistic halting frameworks such as PonderNet (Banino et al., 2021). Building on these foundations, recent work has begun to characterize when additional inference compute actually generalizes beyond training-time budgets (Schwarzschild et al., 2021), how to mitigate failures due to excessive computation (“overthinking”) (Bansal et al., 2022), how to ensure stable dynamics with repeated iteration (Bear et al., 2024), and how looped Transformers can better generalize to out of distribution tasks at test time (McLeish et al., 2024).

Proposition 4.2 (Recurrent attention patterns change slowly under state convergence). Fix a layer ℓ in the recurrent block and consider its attention weight matrices to be tied across recurrences (so W𝑄,ℓ , W𝐾 ,ℓ are the same for all 𝑡). Let ∥·∥ denote a submultiplicative matrix norm that is invariant under transposition (e.g. Spectral or Frobenius norms). Assume the corresponding attention inputs are bounded under this norm as ∥Xℓ,𝑡 ∥ ≤ 𝐵 for all 𝑡. Define 𝜅ℓ = ∥W𝑄,ℓ W𝐾⊤,ℓ ∥. Then, writing Sℓ (X) := softmax( 𝐴ℓ (X)) with 𝐴ℓ (·) as defined above, for any 𝑡 ≥ 1, Sℓ (Xℓ,𝑡 ) − Sℓ (Xℓ,𝑡 −1 )

2𝐵 𝜅ℓ Xℓ,𝑡 − Xℓ,𝑡 −1 , ≤ 𝐿 sm √ 𝑑

where 𝐿 sm is a Lipschitz constant of the row-wise softmax with respect to the chosen norm. Since these attention patterns are characteristic of the different mixing stages of inference (defining, for example, ColSum concentration), we see that mixing behavior will tend towards being constant across recurrences.

4. Looped Transformers Tend Towards the Same Attention Patterns Recent work (Bai et al., 2019; Bansal et al., 2022; Anil et al., 2022) has noted that weight-tied Transformer models often tend towards consistent behavior with repeated iterations. Often1 , this takes the form of convergence to a fixed point X ′ = S 𝑘 (X ′ ). We motivate our work by noting that if this is true for a model with cyclic recurrence, it is also true cyclically:

4.1. Empirical Validation We focus our attention on three different pretrained looped language models: Ouro 1.4B (Zhu et al., 2025), Huginn0125 (Geiping et al., 2025), and Llama with retrofitted recurrence (McLeish et al., 2025). Additional models can be found in App. D, with architecture and training choices summarized in Table 1. These models cover a range of design choices; later in Sec. 4.2 we will isolate the impact of these architectural differences. Except where otherwise specified, all results are visualized on the same random subset of 256 examples from the GSM8k test set; additional results targetting non-reasoning behavior are presented in App. E.4, but we observe no significant changes and the conclusions of the main text remain unchanged.

Proposition 4.1 (Cyclic recurrent blocks reach cyclic fixed points). Let (𝑙, 𝑘)-Recurrent block reach a fixed point X ′ such that S 𝑘 (X ′ ) = X ′ . Then any cyclic permutation of blocks 1, . . . , 𝑘 will also have reached a fixed point. We highlight however that these fixed points are not necessarily the same: the action of each successive layer doesn’t necessarily result in the same point, and the cycle of layers can instead trace out an arbitrary cycle in latent space. Indeed, in Sec. 4.1 we demonstrate that this cyclic fixed point behavior – illustrated in Fig. 1 – is observed frequently in practice. The alternative, where all layers result in the same fixed point, requires that the action of each Transformer block tends to zero.

We start by plotting the norm of the residual stream difference between subsequent iterations in Fig. 3 (and the same for cosine similarities in Fig. 24). This validates the starting assumption of Proposition 4.2 that looped models tend towards behavior in which the layerwise residual stream does not significantly change between recurrences. For each of these models, we plot Frobenius norm between the realized attention matrices at different layers, for 8 loops of the recurrent block in Fig. 2. The diagonal patterns of high similarity demonstrate that the attention matrices of any given layer are most similar to those of the same layer at different recurrences – as predicted by Proposition 4.2.

Convergence to this cyclic behavior implies that the residual stream tends towards being similar across recurrences. Given that block weights are also shared across recurrent iterations – and assuming that the inputs to each block are 1As originally observed by Geiping et al. (2025), other forms of limiting behavior can occur. In App. C we show that 1) this is rare, the vast majority of tokens reach a fixed-point and 2) even with other limiting behavior, stages of inference remain constant.

2 This is a reasonable assumption since all models considered in this work apply a norm before the attention block.

4

A Mechanistic Analysis of Looped Language Models Retrofitted Llama

48 72

10

96

8

120

6

144

4

168

2

8

0

6

7

4

12

Avg. Frobenius Norm

12

Avg. Frobenius Norm

14

24

Huginn-0125

0

6

18

5

24

4

30 36

3

42

2

48

1

10

8

8

12 16

6

20 4

24 28

2

Avg. Frobenius Norm

Ouro 1.4B 0

32

54 0

24 48 72 96 120 144 168

0

6 12 18 24 30 36 42 48 54

Layer Index

0

4

8

Layer Index

12 16 20 24 28 32 Layer Index

Figure 2. Frobenius norm between attention patterns at different depths, averaged across the batch and head dimensions. Depth index visualized on each axis, cells show the norms between attention patterns at each pair of depth indices. Left: Ouro 1.4B (Zhu et al., 2025). Center: Retrofitted Llama (McLeish et al., 2025). Right: Huginn-0125 (Geiping et al., 2025). All models looped 8 times. R.Llama

Late

10

10

40

5

5

20 0

0

0 0

100

PC 2

60

15

Layer

Difference Norm

Recurrences 0-16

Huginn-0125

15

0

Recurrence

100

0

Recurrence

100

We note that this convergence towards similar attention patterns occurs remarkably quickly: for looped Ouro attention patterns appear to converge after the first iteration, and both Huginn-0125 and the retrofitted Llama model demonstrate this cyclic behavior immediately following the prelude. Huginn-0125

60 10 40

10

5

0

20 0

0 0

100 Recurrence

0

100 Recurrence

0

100

0 -1

-2

-2 0

10

-10

0

10

Early

PC 1

Where a model does reach a fixed point, this implies that the action of the entire recurrent block tends towards tracing out a consistent cycle in latent space. We visualize this for retrofitted Llama in Fig. 5.

Layer

Fixed Point Norm

80

15

20

0 -1

see that while Huginn-0125 and Retrofitted Llama demonstrate fast convergence to a fixed point, Ouro does not. As discussed by Bansal et al. (2022); Anil et al. (2022), this supports the suggestion that input injection encourages fixedpoint convergence: we investigate this further in Sec. 4.2.

Late 30

1

Figure 5. Retrofitted Llama (McLeish et al., 2025) latent space trajectory traced out by the hidden states of the final sequence position on a single test prompt; reduced to two dimensions by computing PCA over all final sequence position embeddings. Trajectories perfectly overlap in the second plot, demonstrating that a cyclic fixed point has been reached.

Recurrence

R.Llama

1

Late

PC 1

Figure 3. Norm of the difference between the residual stream after successive recurrences of the same layer.

Ouro 1.4B

2

-10

Early

Recurrences 16-32

2

Recurrence

Ouro 1.4B

Early

4.2. Impact of Architecture Choices

Recurrence

Figure 4. Norm of the difference between the residual stream after each layer in the recurrent block and its “approximate fixed point” - the residual stream after that layer in the 128th recurrence. While Huginn-0125 and retrofitted Llama quickly reach a fixed point, Ouro does not - despite small successive differences evidenced by Fig. 3.

Several existing works (Bansal et al., 2022; Anil et al., 2022) have noted that input injection is important in order for a recurrent model to reach a fixed point: in this section we replicate this finding and supplement with additional insights on the impact of norm structure in reaching a fixed point. We conduct a series of experiments on randomly initialized models. These demonstrate similar cyclic behavior to their trained counterparts, suggesting that behavior observed here is likely to generalize to the cyclic behavior of trained models. We compare pre-norm (used by the retrofitted recurrent models) and the norms used by the Huginn-0125 and Ouro models, testing each both with and without input injection: in this way we test the most significant architectural differences between the pretrained Looped models tested;

However, despite the tendency visible in Fig. 3 towards small changes between successive iterations we discover that it is not the case that looped models always reach a fixed point, or even consistent limiting behavior: for each model we find an “approximate fixed point” per layer by iterating 128 times, then compute the norm of the difference between the output of each layer at every recurrence and its corresponding fixed point. This is visualized in Fig. 4. We 5

A Mechanistic Analysis of Looped Language Models Pre

Huginn-0125

Ouro

Input Inj.

of Retrofitted Llama in Fig. 7, revealing consistent mixing cycles that repeat with every iteration of the recurrent block. However, each individual layer (solid colorful lines) changes very little in realized depth: after an initial transitory phase they quickly converge towards constant behavior.

Lowest Similarity Fixed Point 1.0

1.00

0.8 0.75

0.6

0.50

0.4

0.25

0.2

0.00

0.0 -0.2

-0.25 0

20

40

Recurrence

60

True Layer Order L4

0.7 Colsum Concentration

Fixed Point Cosine Sim

Same Layer

0

20

40

60

Recurrence

Figure 6. Cosine similarity between residual streams after each layer and the approximate fixed point for a range of norms, with and without input injection. Each model is randomly initialized with 12 layers. Cosine similarity is taken between the residual stream after the first layer at each recurrence and left: the approximate fixed point of the first layer, right: the approximate fixed point of the layer with the lowest cosine similarity to the first layer.

L5 L6

L7 L8

L9

0.6 0.5 0.4 0.3 0

10

20

30

40

50

Recurrent Position

Figure 7. Stages of inference for retrofitted Llama (McLeish et al., 2025) with 8 recurrences. Individual layers are visualized as solid lines; successive layers in the looped Transformer as a dashed black line. Individual layers quickly converge towards constant behavior, and the cyclic action of these layers results in cyclic stages of inference.

see Table 1 for details. Each model has 12 layers with no prelude or coda; see Fig. 31 for alternative configurations. Our results are visualized in Fig. 6, where we visualise the mean over 3 random model initializations for each configuration. We see that input injection results in stable fixed point behavior for all norm types other than Ouro, whereas omitting input injection means that only pre-norm reaches a stable fixed point. However, this fixed point reached by pre-norm without input injection is a “degenerate” one: each layer converges to the same fixed point. This can be determined from the rightmost frame of Fig. 6, which demonstrates that the lowest cosine similarity between the first layer and any other layer’s fixed point still converges to 1.

Instead of occurring throughout the realized depth of the looped model, we find that the familiar feedforward stages of inference occur within each looped block. Fig. 8 demonstrates that ColSum concentration within each looped block closely resembles that of feedforward models. We draw attention to two observations: 1) Ouro 1.4B, despite being trained from scratch with recurrence, mirrors Llama mixing stages and 2) the retrofitted models closely follow the stages of inference of their associated base model, but the initial and final stages are performed only once by the prelude and coda respectively, while the “middle” stages are repeated in the recurrent block. It is particularly remarkable that these stages of inference appear in each Ouro recurrent block when pretraining from initialization; we discuss further the formation of stages of inference, and attempt to isolate their formation from specific training procedures, in Sec. 5.1.

5. Stages of Inference in Looped Models Mirror Feedforward Computation The previous section shows that, empirically, a wide range of models converge to a regime in which attention patterns within individual layers change only minimally across recurrences. As a result, attention dynamics in looped Transformers are constrained in depth, since layers are cyclically weight-tied to earlier ones. This behavior contrasts with feedforward Transformers, which impose no such constraints and exhibit sharp, layer-wise changes in attention patterns across depth. Prior work has linked these sharp transitions to characteristic stages of inference (Lad et al., 2024; Queipo-de Llano et al., 2025), introduced in Sec. 2.2. In this section, we study how cyclic weight sharing alters these stages of inference in looped Transformers. In the main text we frame our analysis using ColSum concentration, a metric for identifying stages of inference introduced in Sec. 2.2, with extensive additional results in App. E.3.

However, Huginn-0125 does not demonstrate clear stages of inference (Fig. 36). We suggest that this is likely due to the specific norm structures used by these models (Table 1). Huginn-0125 and Ouro both use a “sandwich” norm structure, but Huginn-0125 implements this by normalizing the residual streams whereas Ouro instead normalizes the outputs of the attention and MLP units, only normalizing the residual stream at the end of each recurrent block. We demonstrate the impact of the different norms by plotting residual stream magnitudes for a range of models in Fig. 9. This means that Huginn-0125 is unable to develop the growth in residual stream magnitude that Queipo-de Llano et al. (2025, Section 3.3) note causes compression behavior, leading to sink formation and stages of inference.

We visualize ColSum concentration over the realized depth 6

A Mechanistic Analysis of Looped Language Models

Colsum Concentration

Retrofitted Llama (8 steps) 0.7

Llama 3.2 1B

0.6

0.6

Retrofitted OLMo (8 steps)

0.4 0.3

0.4

0.2

0.3 0.2

0.1

0.2 0

20

40

60

80

100

Prelude Coda OLMo2

0.5

0.5 0.4

Late

0.6

Prelude Coda Llama 3.2 1B

Recurrence

Ouro 1.4B (4 steps) 0.8

0

% Block Depth

20

40

60

80

% Block Depth

100

0

20

40

60

80

100

Early

% Block Depth

Figure 8. Stages of inference for each recurrent loop in left: Ouro 1.4B (Zhu et al., 2025) center: retrofitted Llama and right: retrofitted OLMo (McLeish et al., 2025). Ouro 1.4B resembles Llama stages of inference, and the two retrofitted to their associated base models.

(2, 12 ⊗ 4, 2) are plotted in Fig. 10. We compare these to a “control” feedforward network without recurrence, of depth 12. These experiments are on a small scale and as such need to be treated with caution. However, they appear to provide initial evidence that – even without training methods that may bias towards feedforward stages of inference – looped models have a tendency to self-organize into multiple different mixing stages in recurrent depth, which resemble feedforward stages. The fact that the optimization of these models results in these mixing stages of inference suggests that they are beneficial to language modeling even when applied repeatedly in recurrent depth. We additionally test the impact of input injection and sandwich layers in App. E.6.

Residual Norm

103

102 Huginn-0125 Ouro 1.4B Retrofitted Llama

101

0

20

40

60

80

100

% Recurrent Position

Figure 9. Norms of the residual stream for a range of models, demonstrating that Huginn-0125 is unable to develop the activation magnitude changes required for stages of inference due to its repeated normalization of the residual stream.

5.2. Stability to Unseen Numbers of Recurrences 5.1. Self-Organization Into Stages of Inference

Sec. 4 establishes that some looped Transformer architectures converge to a fixed point (such as the retrofitted series and Huginn-0125) while others do not (such as Ouro). However, when evaluating these looped Transformers within the range of recurrences on which they were trained (Sec. 5), we see extremely similar behavior: convergence to feedforward stages of inference within each recurrent block.

An open question from our analysis of pretrained models (Fig. 8) is whether stages of inference emerge naturally during pretraining, or whether they are instead induced by specific aspects of the training procedure. Several mechanisms may introduce an implicit bias toward feedforward stages of inference: retrofitted models (McLeish et al., 2025) may inherit the inference stages of the underlying base model; models trained with recurrence schedulers that permit a single recurrence (Geiping et al., 2025; McLeish et al., 2025) are partially optimized as feedforward models; training objectives that decompose into separate loss terms for each recurrence (Zhu et al., 2025) partially correspond to feedforward model training.

In this section, we demonstrate that a significant difference arises when generalizing to unseen test-time recurrence depths: models that do not reach a fixed point exhibit unstable stages of inference. Conversely, models with recurrent-block-wise stages of inference that also converge to a fixed point are guaranteed to keep enacting these stages of inference for arbitrary test time recurrences.

Therefore in this section we seek to investigate whether stages of inference can arise in Looped Transformers without these training biases. To explore this we pre-train several small-scale Looped Transformers, explicitly removing the biases above: we pre-train from scratch with a constant recurrence of 4, using a standard loss that considers only the final latent state when predicting the next token. Our code and training procedure are adapted from Karpathy (2025); additional details can be found in App. B.

We first verify this stability in Fig. 11, an extended plot of Fig. 7. This demonstrates that for retrofitted Llama (results hold for other models using input injection), each individual layer quickly reaches a stable states and then exhibits consistent stages of inference behavior for an arbitrary number of test time recurrences. However, Ouro 1.4B does not exhibit this behavior, with individual layers changing continuously throughout later recurrences. The effect that this has on stages of inference is visualized in Fig. 12.

ColSum concentrations for these trained Looped Transformers with configurations (2, 4 ⊗ 4, 2), (2, 8 ⊗ 4, 2) and

Existing research implies that this stability correlates with 7

0.4 0.3

0.4 0.4 0.3

0.2

0.3

0.2

0.1

0.2

0.1 0.0

0.2

0.4

Late

0.5

Prelude Coda Feedforward

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

Recurrence

Colsum Concentration

A Mechanistic Analysis of Looped Language Models

0.1 0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

0.0

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 10. ColSum concentrations for small-scale trained Looped Transformers with a simplified loss function and constant train recurrence schedule of 4 recurrences. Also visualized in red is a “control” feedforward Transformer of depth 12. All models have 2 prelude and 2 coda layers with no input injection, left: 4 recurrent layers, center: 8 recurrent layers, right: 12 recurrent layers. Ouro 1.4B (128 steps)

0.5

across a range of architectures that recurrent blocks tend to “mirror” the stages of a feedforward Transformer, and provide evidence that this may be emergent behavior learned during training, even when not explicitly encouraged by the training process. We further investigate the implications for these mixing stages when models converge to a stable fixed point, and when they do not.

Late

0.8 0.6

0.4

0.4

0.3

0.2 0

50

100

Layer

Colsum Concentration

Retrofitted Llama (128 steps)

0

Recurrence

50

Early

100

Recurrence

Implications of Findings The implications of our findings are bidirectional. On the one hand, the structure of looped architectures provides a novel lens to study stages of inference, while tracking these stages simultaneously reveals the internal mechanics of recurrent depth. In particular, since looped models decouple functional depth from parameter count, they provide an interesting new perspective on why these stages of inference form: previous work had suggested that these stages exist to mitigate the “harms” of transformer depth (Barbero et al., 2025; Queipo-de Llano et al., 2025), however our work shows that looped models develop these same stages while simultaneously improving performance with greater recurrent depth. On the other hand, our findings that looped models exhibit (and selforganize into) similar stages of inference to feedforward models means that insights from the feedforward setting can be applied to looped models: predictable stages offer actionable pathways for efficient architectural design, including stage-dependent attention sparsification and the leaner parameterization of middle-stage MLPs where representations are reliably compressed and low-rank.

Figure 11. Colsum concentration of each layer with successive recurrences for left: retrofitted Llama and right: Ouro 1.4B, both using 128 recurrences. While the layers of retrofitted Llama quickly converge to constant ColSum concentration, the layers of Ouro continually change throughout the recurrences tested. Ouro 1.4B (128 steps) Late

0.8

0.7 0.6

Recurrence

Colsum Concentration

Retrofitted Llama (128 steps)

0.6

0.5 0.4

0.4 0.3

0.2

0.2 0

50 % Block Depth

100

0

50

100

Early

% Block Depth

Figure 12. Colsum concentration of each layer vs the percentage depth at which that layer appears in the recurrent block. Left: retrofitted Llama and right: Ouro 1.4B, both using 128 recurrences. Feedforward Llama shown in dashed red.

out-of-domain performance. Models which exhibit “stable” stages of inference for arbitrary test time iterations also avoid performance deterioration in this extrapolation regime: whereas extrapolation beyond training recurrences harms the performance of Ouro (Zhu et al., 2025, Tab. 10), Huginn-0125 performance remains constant in this extrapolation region (Geiping et al., 2025, Fig. 1).

Limitations and Future Work We focus exclusively on cyclic recurrence as this appears to be the dominant approach in the literature. However, this means our analysis does not extend to sequential recurrence with multiple separate recurrent blocks; for analysis of this setting we refer the reader to Pappone et al. (2025). Despite empirically investigating the architectural choices that result in stable limiting behavior in looped models, we have not established analytically why this is the case, nor whether this stable limiting behavior is desirable or restrictive for reasoning tasks.

6. Conclusion This paper examines the limiting behavior of Looped Transformers, exploring implications for “mixing” stages of inference observed in feedforward models. We demonstrate 8

A Mechanistic Analysis of Looped Language Models

Impact Statement

Banino, A., Balaguer, J., and Blundell, C. Pondernet: Learning to ponder. arXiv preprint arXiv:2107.05407, 2021.

Ethical aspects and future societal consequences of this particular work are limited. The goal of our work is to advance understanding of looped Language Models, which themselves seem to demonstrate strong reasoning performance; as such, our work is in support of a field that has potential societal consequences if future models are able to undertake more advanced reasoning tasks. However, we do not within this work introduce any more powerful reasoning models, and the impact of our work is limited to understanding existing models, and potentially guiding future advancements.

Bansal, A., Schwarzschild, A., Borgnia, E., Emam, Z., Huang, F., Goldblum, M., and Goldstein, T. End-toend algorithm synthesis with recurrent networks: Logical extrapolation without overthinking. Advances in Neural Information Processing Systems, 35:20232–20242, 2022. Barbero, F., Banino, A., Kapturowski, S., Kumaran, D., Madeira Araújo, J., Vitvitskyi, O., Pascanu, R., and Veličković, P. Transformers need glasses! information over-squashing in language tasks. Advances in Neural Information Processing Systems, 37:98111–98142, 2024.

Acknowledgments

Barbero, F., Arroyo, A., Gu, X., Perivolaropoulos, C., Bronstein, M., Veličković, P., and Pascanu, R. Why do llms attend to the first token? arXiv preprint arXiv:2504.02732, 2025.

HB acknowledges funding support from the EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines and Systems No. EP/S024050/1. MB is partially supported by the EPSRC Turing AI World-Leading Research Fellowship No. EP/X040062/1 and EPSRC AI Hub No. EP/Y028872/1.

Bear, J., Prugel-Bennett, A., and Hare, J. Rethinking deep thinking: Stable learning of algorithms using lipschitz constraints. Advances in Neural Information Processing Systems, 37:97027–97052, 2024.

References

Blayney, H., Arroyo, Á., Dong, X., and Bronstein, M. M. glstm: Mitigating over-squashing by increasing storage capacity. arXiv preprint arXiv:2510.08450, 2025.

Alon, U. and Yahav, E. On the bottleneck of graph neural networks and its practical implications. arXiv preprint arXiv:2006.05205, 2020.

Cai, C. and Wang, Y. A note on over-smoothing for graph neural networks. arXiv preprint arXiv:2006.13318, 2020.

Anil, C., Pokle, A., Liang, K., Treutlein, J., Wu, Y., Bai, S., Kolter, J. Z., and Grosse, R. B. Path independent equilibrium models can better exploit test-time computation. Advances in Neural Information Processing Systems, 35: 7796–7809, 2022.

Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.

Arroyo, Á., Gravina, A., Gutteridge, B., Barbero, F., Gallicchio, C., Dong, X., Bronstein, M., and Vandergheynst, P. On vanishing gradients, over-smoothing, and oversquashing in gnns: Bridging recurrent and graph learning. arXiv preprint arXiv:2502.10818, 2025.

Csordás, R., Irie, K., Schmidhuber, J., Potts, C., and Manning, C. D. Moeut: Mixture-of-experts universal transformers. Advances in Neural Information Processing Systems, 37:28589–28614, 2024.

Arroyo, A., Barbero, F., Blayney, H., Bronstein, M. M., Dong, X., Lio, P., Pascanu, R., and Vandergheynst, P. A survey on over-smoothing and over-squashing: Unified propagation perspectives on graph neural networks and transformers. Transactions on Machine Learning Research, 2026. ISSN 2835-8856.

Darlow, L., Regan, C., Risi, S., Seely, J., and Jones, L. Continuous thought machines. arXiv preprint arXiv:2505.05522, 2025. Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł. Universal transformers. arXiv preprint arXiv:1807.03819, 2018.

Bae, S., Kim, Y., Bayat, R., Kim, S., Ha, J., Schuster, T., Fisch, A., Harutyunyan, H., Ji, Z., Courville, A., et al. Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation. arXiv preprint arXiv:2507.10524, 2025.

Geiping, J., McLeish, S., Jain, N., Kirchenbauer, J., Singh, S., Bartoldson, B. R., Kailkhura, B., Bhatele, A., and Goldstein, T. Scaling up test-time compute with latent reasoning: A recurrent depth approach. arXiv preprint arXiv:2502.05171, 2025.

Bai, S., Kolter, J. Z., and Koltun, V. Deep equilibrium models. Advances in neural information processing systems, 32, 2019.

Graves, A. Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983, 2016. 9

A Mechanistic Analysis of Looped Language Models

Nair, P. Softmax is 1/2-lipschitz: A tight bound across all ℓ 𝑝 norms. arXiv preprint arXiv:2510.23012, 2025.

Gu, X., Pang, T., Du, C., Liu, Q., Zhang, F., Du, C., Wang, Y., and Lin, M. When attention sink emerges in language models: An empirical view. arXiv preprint arXiv:2410.10781, 2024.

Pappone, F., Crisostomi, D., and Rodolà, E. Two-scale latent dynamics for recurrent-depth transformers. arXiv preprint arXiv:2509.23314, 2025.

Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.

Queipo-de Llano, E., Arroyo, Á., Barbero, F., Dong, X., Bronstein, M., LeCun, Y., and Shwartz-Ziv, R. Attention sinks and compression valleys in llms are two sides of the same coin. arXiv preprint arXiv:2510.06477, 2025.

Gurnee, W., Horsley, T., Guo, Z. C., Kheirkhah, T. R., Sun, Q., Hathaway, W., Nanda, N., and Bertsimas, D. Universal neurons in GPT2 language models. Transactions on Machine Learning Research, 2024. ISSN 2835-8856.

Sandoval-Segura, P., Wang, X., Panda, A., Goldblum, M., Basri, R., Goldstein, T., and Jacobs, D. Using attention sinks to identify and evaluate dormant heads in pretrained llms. arXiv preprint arXiv:2504.03889, 2025.

Hariri, A., Arroyo, Á., Gravina, A., Eliasof, M., Schönlieb, C.-B., Bacciu, D., Azizzadenesheli, K., Dong, X., and Vandergheynst, P. Return of chebnet: Understanding and improving an overlooked gnn on long range tasks. arXiv preprint arXiv:2506.07624, 2025.

Saunshi, N., Dikkala, N., Li, Z., Kumar, S., and Reddi, S. J. Reasoning with latent thoughts: On the power of looped transformers. arXiv preprint arXiv:2502.17416, 2025. Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., and Metzler, D. Confident adaptive language modeling. Advances in Neural Information Processing Systems, 35:17456–17472, 2022.

Jolicoeur-Martineau, A. Less is more: Recursive reasoning with tiny networks. arXiv preprint arXiv:2510.04871, 2025. Karpathy, A. nanochat: The best ChatGPT that $100 can buy, 2025. URL https://github.com/karpathy/ nanochat.

Schwarzschild, A., Borgnia, E., Gupta, A., Huang, F., Vishkin, U., Goldblum, M., and Goldstein, T. Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks. Advances in Neural Information Processing Systems, 34:6695–6706, 2021.

Ke, Y., Li, X., Liang, Y., Shi, Z., and Song, Z. Advancing the understanding of fixed point iterations in deep neural networks: A detailed analytical study. arXiv preprint arXiv:2410.11279, 2024.

Skean, O., Arefin, M. R., Zhao, D., Patel, N. N., Naghiyev, J., LeCun, Y., and Shwartz-Ziv, R. Layer by layer: Uncovering hidden representations in language models. In Fortysecond International Conference on Machine Learning, 2025.

Koishekenov, Y., Lipani, A., and Cancedda, N. Encode, think, decode: Scaling test-time reasoning with recursive latent thoughts. arXiv preprint arXiv:2510.07358, 2025.

Sun, M., Chen, X., Kolter, J. Z., and Liu, Z. Massive activations in large language models. In First Conference on Language Modeling, 2024.

Lad, V., Lee, J. H., Gurnee, W., and Tegmark, M. The remarkable robustness of llms: Stages of inference? arXiv preprint arXiv:2406.19384, 2024.

Tan, S., Shen, Y., Chen, Z., Courville, A., and Gan, C. Sparse universal transformer. arXiv preprint arXiv:2310.07096, 2023.

Lu, W., Yang, Y., Lee, K., Li, Y., and Liu, E. Latent chainof-thought? decoding the depth-recurrent transformer. arXiv preprint arXiv:2507.02199, 2025.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.

McLeish, S., Bansal, A., Stein, A., Jain, N., Kirchenbauer, J., Bartoldson, B., Kailkhura, B., Bhatele, A., Geiping, J., Schwarzschild, A., et al. Transformers can do arithmetic with the right embeddings. Advances in Neural Information Processing Systems, 37:108012–108041, 2024.

Veličković, P., Perivolaropoulos, C., Barbero, F., and Pascanu, R. Softmax is not enough (for sharp size generalisation). arXiv preprint arXiv:2410.01104, 2024.

McLeish, S., Li, A., Kirchenbauer, J., Kalra, D. S., Bartoldson, B. R., Kailkhura, B., Schwarzschild, A., Geiping, J., Goldstein, T., and Goldblum, M. Teaching pretrained language models to think deeper with retrofitted recurrence. arXiv preprint arXiv:2511.07384, 2025.

Wang, G., Li, J., Sun, Y., Chen, X., Liu, C., Wu, Y., Lu, M., Song, S., and Yadkori, Y. A. Hierarchical reasoning model. arXiv preprint arXiv:2506.21734, 2025. 10

A Mechanistic Analysis of Looped Language Models

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453, 2023. Xu, K. and Sato, I. On expressive power of looped transformers: Theoretical analysis and enhancement via timestep encoding. arXiv preprint arXiv:2410.01405, 2024. Yang, L., Lee, K., Nowak, R., and Papailiopoulos, D. Looped transformers are better at learning learning algorithms. arXiv preprint arXiv:2311.12424, 2023. Yudin, N., Gaponov, A., Kudriashov, S., and Rakhuba, M. Pay attention to attention distribution: A new local lipschitz bound for transformers. arXiv preprint arXiv:2507.07814, 2025. Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. Zhu, R.-J., Wang, Z., Hua, K., Zhang, T., Li, Z., Que, H., Wei, B., Wen, Z., Yin, F., Xing, H., et al. Scaling latent reasoning via looped language models. arXiv preprint arXiv:2510.25741, 2025.

11

A Mechanistic Analysis of Looped Language Models

Appendix Contents A Proofs of Propositions

13

B Additional Experimental Details

14

C Non-Fixed-Point Limiting Behavior

15

C.1 How Frequent is Non-Fixed-Point Behavior? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

15

C.2 Do Intermediate Layers Exhibit Non-Fixed-Point Behavior? . . . . . . . . . . . . . . . . . . . . . . . .

19

C.3 How Does Non-Fixed-Point Behavior Impact Stages of Inference? . . . . . . . . . . . . . . . . . . . . .

21

D Additional Fixed Point Results

22

D.1 Cyclic Similarity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

22

D.2 Fixed Point and Successive Differences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

23

D.3 Latent Space Trajectories . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

25

D.4 Architecture Choices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26

E Additional Stages of Inference Results

27

E.1 Input Independent Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

27

E.2 Input Dependent Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

28

E.3 Cyclic Stages of Inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

28

E.4 Non-Reasoning Stages of Inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

32

E.5 Stability To Unseen Test-Time Recurrences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

34

E.6 How Architecture Choices Affect the Formation of Stages of Inference . . . . . . . . . . . . . . . . . . .

36

F Looped Floorplan

39

12

A Mechanistic Analysis of Looped Language Models

A. Proofs of Propositions Proof of Proposition 4.1 Proof sketch. Note S 𝑘 (X ′ ) = B 𝑘 (B 𝑘−1 (. . . B1 (X ′ ) . . . )) = X ′ Applying B1 to both sides yields to B1 (B 𝑘 (B 𝑘−1 (. . . B1 ( X ′ ) . . . ))) = B1 (X ′ ). Defining Y ′ = B1 (X ′ ), we obtain B1 (B 𝑘 (B 𝑘−1 (. . . B2 (Y ′ ) . . . ))) = Y ′ . The general proof follows by induction, and extends trivially to input injection. □ Proof. Assume that (𝑙, 𝑘)-Recurrent block 𝑆 𝑘 reaches a fixed point such that 𝑆 𝑘 (X ′ ) = X ′ . Note that S 𝑘 (X ′ ) = B 𝑘−1 (B 𝑘−2 (. . . B0 (X ′ ) . . . )) = X ′ define the cyclic shift function 𝑓𝑛 (𝑖) = (𝑖 + 𝑛) mod 𝑘. Now we aim to prove by induction ∀𝑛 ∈ Z+ ∪ {0} : B 𝑓𝑛 (𝑘−1) (B 𝑓𝑛 (𝑘−2) (. . . B 𝑓𝑛 (0) (Z ′ ) . . . )) = Z ′ for some Z ′ . The base case 𝑛 = 0 is trivial as ∀𝑖 < 𝑘 : 𝑓0 (𝑖) = 𝑖. Now assume for some 𝑛 = 𝑗 B 𝑓 𝑗 (𝑘−1) (B 𝑓 𝑗 (𝑘−2) (. . . B 𝑓 𝑗 (0) (Z ′ ) . . . )) = Z ′

(8)

Now, let 𝑛 = 𝑗 + 1. Note 𝑓 𝑗+1 (𝑖) = (𝑖 + 𝑗 + 1)

mod 𝑘 = 𝑓 𝑗 (𝑖 + 1)

Therefore B 𝑓 𝑗+1 (𝑘−1) (B 𝑓 𝑗+1 (𝑘−2) (. . . B 𝑓 𝑗+1 (0) (Y ) . . . )) = B 𝑓 𝑗 (𝑘 ) (B 𝑓 𝑗 (𝑘−1) (. . . B 𝑓 𝑗 (1) (Y ) . . . ))

(9)

= B 𝑓 𝑗 (0) (B 𝑓 𝑗 (𝑘−1) (. . . B 𝑓 𝑗 (1) (Y ) . . . ))

(10)

Now take Eq. (8) and apply B 𝑓 𝑗 (0) to both sides, defining a new fixed point Z ′′ = B 𝑓 𝑗 (0) (Z ′ ): B 𝑓 𝑗 (0) (B 𝑓 𝑗 (𝑘−1) (B 𝑓 𝑗 (𝑘−2) (. . . B 𝑓 𝑗 (0) (Z ′ ) . . . ))) = B 𝑓 𝑗 (0) (Z ′ ) ′′

B 𝑓 𝑗 (0) (B 𝑓 𝑗 (𝑘−1) (B 𝑓 𝑗 (𝑘−2) (. . . B 𝑓 𝑗 (1) (Z ) . . . ))) = Z

′′

(11) (12)

Combining Eq. (12) and Eq. (10), we see that there exists a fixed point Z ′′ such that B 𝑓 𝑗+1 (𝑘−1) (B 𝑓 𝑗+1 (𝑘−2) (. . . B 𝑓 𝑗+1 (0) (Z ′′ ) . . . )) = Z ′′ □

Therefore completing the induction step and proving the proposition. Proof of Proposition 4.2 Proof. Define  Sℓ (X) := softmax( 𝐴ℓ (X)) = softmax

XW𝑄 W𝐾⊤ X ⊤ √ 𝑑



Let 𝐿 sm be the Lipschitz constant of the row-wise softmax such that Sℓ (Xℓ,𝑡 ) − Sℓ (Xℓ,𝑡 −1 ) ≤ 𝐿 sm 𝐴ℓ (Xℓ,𝑡 ) − 𝐴ℓ (Xℓ,𝑡 −1 ) Recently Nair (2025) shows that 𝐿 sm = 1/2. 13

(13)

A Mechanistic Analysis of Looped Language Models

Define for convenience M = W𝑄 W𝐾⊤ and Δℓ,𝑡 = Xℓ,𝑡 − Xℓ,𝑡 −1 . Then ⊤ ⊤ Xℓ,𝑡 −1 W𝑄 W𝐾⊤ Xℓ,𝑡 Xℓ,𝑡 W𝑄 W𝐾⊤ Xℓ,𝑡 −1 − √ √ 𝑑 𝑑   ⊤ 1  ⊤ =√ Δℓ,𝑡 + Xℓ,𝑡 −1 M Δℓ,𝑡 + Xℓ,𝑡 −1 − Xℓ,𝑡 −1 M Xℓ,𝑡 −1 𝑑  1  ( ( ⊤ ⊤ ⊤ ⊤( ⊤( ( ( ( ( = √ Δℓ,𝑡 M Δℓ,𝑡 + Δℓ,𝑡 M Xℓ,𝑡 + X M Δ + X M X − X M X ( ( ℓ,𝑡 −1 ℓ,𝑡 −1 ℓ,𝑡 −1 −1 ℓ,𝑡 (( ℓ,𝑡 −1 (( ℓ,𝑡 −1 𝑑   ⊤ 1 ⊤ = √ Δℓ,𝑡 M Δℓ,𝑡 + Xℓ,𝑡 −1 + Xℓ,𝑡 −1 M Δℓ,𝑡 𝑑  1  ⊤ ⊤ + Xℓ,𝑡 −1 M Δℓ,𝑡 = √ Δℓ,𝑡 M Xℓ,𝑡 𝑑

𝐴ℓ (Xℓ,𝑡 ) − 𝐴ℓ (Xℓ,𝑡 −1 ) =

Therefore,  1  ⊤ ⊤ 𝐴ℓ (Xℓ,𝑡 ) − 𝐴ℓ (Xℓ,𝑡 −1 ) = √ Δℓ,𝑡 M Xℓ,𝑡 + Xℓ,𝑡 −1 M Δℓ,𝑡 𝑑  1  ⊤ ⊤ ≤ √ Δℓ,𝑡 M Xℓ,𝑡 + Xℓ,𝑡 −1 M Δℓ,𝑡 𝑑  M Xℓ,𝑡 + Xℓ,𝑡 −1 ≤ Δℓ,𝑡 √ 𝑑 Assuming ∥Xℓ,𝑡 ∥ ≤ 𝐵 for all 𝑡, and defining 𝜅ℓ = ∥W𝑄,ℓ W𝐾⊤,ℓ ∥ = M , we see 2𝐵𝜅ℓ Xℓ,𝑡 − Xℓ,𝑡 −1 𝐴ℓ (Xℓ,𝑡 ) − 𝐴ℓ (Xℓ,𝑡 −1 ) ≤ √ 𝑑

(14)

Combining Eq. (13) and Eq. (14), we complete the proof: 2𝐵𝜅ℓ Sℓ (Xℓ,𝑡 ) − Sℓ (Xℓ,𝑡 −1 ) ≤ 𝐿 sm √ Xℓ,𝑡 − Xℓ,𝑡 −1 𝑑

(15) □

B. Additional Experimental Details Unless stated otherwise, all of our experiments are averaged over the same subset of 256 random examples from the test split of the GSM8k (Cobbe et al., 2021) dataset. A few illustrative plots (for example, latent space trajectories) are instead produced with a test sequence that we obtain from Barbero et al. (2025): “Hello! I’ve been well. I hope that you’re doing well.” Additional results targetting non-reasoning behavior using the HellaSwag dataset (following an identical setup of running inference on the same 256 random examples from the test split) can be found in App. E.4. All pretrained models are obtained from Huggingface, model references provided in Table 2. We use standard settings for the tokenizers of each model, and as such some models prepend a BOS token whereas others do not: we make this clear in the ‘Prepends BOS’ column of the same table. Our small training runs in Sec. 5.1 are performed by adapting a publicly available fork of Nanochat (Karpathy, 2025), https://github.com/TrelisResearch/nanochat/tree/recursive. For all experiments we use a model dimension (residual stream) of 512 (4 heads of dimension 128) and train for 3.7B tokens. As discussed in the main text, loss is the same as that of a regular feedforward model: cross entropy loss on the final output representation (as opposed to the summed loss of Zhu et al. (2025)). Each model is trained for a constant 4 recurrences (as opposed to the Poisson sampling of Geiping et al. (2025)). All models use pre-norm only: this norm does not result in the stability issues reported by Geiping et al. (2025); Zhu et al. (2025), but this is likely because we are operating at a far smaller scale. 14

A Mechanistic Analysis of Looped Language Models Table 1. Looped Transformer architecture summary. Retrofitted Llama and OLMo use 6 layers in their recurrent block, TinyLlama uses 8.

Model Name

Structure

Train Loops

Block Norm

Norm After Loop?

Ouro 1.4B (Zhu et al., 2025)

(0, 24, 0)

4

X̂ = X + 𝑛 (Attn(𝑛(X))   X ′ = X̂ + 𝑛 MLP(𝑛( X̂))

Yes

Huginn-0125 et al., 2025)

(2, 4, 2) 𝐼

32

X̂ = 𝑛 (X + Attn(𝑛(X))   X ′ = 𝑛 X̂ + MLP(𝑛( X̂))

Yes

(4, 6, 4) 𝐼 / (4, 8, 4) 𝐼

32

(Geiping

Retrofitted (McLeish et al., 2025) Model Ouro 1.4B

Num Params 1.4B

Ouro 2.6B

2.6B

Huginn-0125

3.5B

Retrofitted Llama

1B

Retrofitted OLMo-2

1B

Retrofitted TinyLlama

0.8B

X̂ = X + Attn(𝑛(X) X ′ = X̂ + MLP(𝑛( X̂))

No

Huggingface ID ByteDance/ Ouro-1.4B ByteDance/ Ouro-2.6B

Base Model -

Prepends BOS ×

Notes -

ByteDance/ Ouro-1.4B

×

tomg-groupumd/huginn0125 smcleish/ RecurrentLlama3.2-trainrecurrence-32 smcleish/ RecurrentOLMo-20425-trainrecurrence-32 smcleish/ RecurrentTinyLlama3T-trainrecurrence-32

-

meta-llama/ Llama-3.2-1B

allenai/OLMo2-0425-1B

×

Models trained with fewer recurrences exhibit similar mixing patterns. As above.

TinyLlama/ TinyLlama1.1Bintermediatestep-1431k-3T

As above.

“Upcycled” from Ouro 1.4B by repeating the same layers after the first training phase. -

Table 2. Additional Huggingface details on pretrained Looped models used.

C. Non-Fixed-Point Limiting Behavior C.1. How Frequent is Non-Fixed-Point Behavior? In this section we investigate more closely the “orbits” and “sliders” initially observed by Geiping et al. (2025). These are important as they appear to represent stable limiting behavior that are not fixed points. We develop a heuristic algorithm to detect these behaviors, presented in Algorithm 1. We use the sequence of cosine similarities from the final recurrent layer (per token), where the output residual stream after each of 128 recurrences is compared to the final residual stream (as in Fig. 22). This is visualized in the leftmost column of Fig. 13. We set threshold 15

A Mechanistic Analysis of Looped Language Models

𝜏 = 0.05 and fixed-point fraction 𝜌 = 0.9. Using this algorithm, we classify the limiting behavior over all tokens in the GSM8k test set for the Huginn-0125 and Retrofitted Llama models. We discover that the system prompt used before presenting the GSM8k question has a large impact on the limiting behavior types and as such we test the following prompts across both models: Long Persona This is the system prompt used by Geiping et al. (2025) to produce their Orbit plots: You are Huginn, an AI assistant who embodies careful thought and deliberation. Your responses demonstrate: Methodical reasoning, breaking complex problems into clear steps Mathematical and programming expertise grounded in fundamentals The ability to acknowledge uncertainty and correct course when needed Clear communication that illuminates rather than just informs When engaging with questions, you first seek to understand their deeper structure before answering. Like your namesake who flew the nine worlds seeking wisdom, you explore problems from multiple angles, helping users build genuine understanding rather than providing shallow answers. You express warmth and intellectual curiosity while maintaining professionalism. When faced with errors or confusion, you model honest reflection and careful correction. Your goal is not just to provide answers, but to help humans develop clearer, deeper thinking. Long Persona (Padded) This is the same prompt as above but with all tokens replaced with the padding token: we include this to test whether the behavior arises purely due to the length of the input. Short Math

This is the system prompt used by Geiping et al. (2025) in their GSM8k benchmark evaluation:

You are a helpful assistant that is capable of helping users with mathematical reasoning. We present the percentage of GSM8k tokens that exhibit each classification of limiting behavior in Table 3 and the percentage of GSM8k examples that exhibit each classification of limiting behavior across any of their tokens in Table 4. These results reveal that these non-fixed-point limiting behaviors appear to be extremely rare in practice: without a system prompt (the setting used throughout this paper) only approximately 0.02% of tokens exhibit non-fixed-point behavior. This percentage can be significantly increased with the longer system prompt, but these behaviors remain rare at 0.14%. Curiously, the “Persona” system prompt used by Geiping et al. (2025) seems to increase the occurrence of these orbits and sliders more than a comparable prompt of padding tokens, suggesting that the effect is not purely to do with sequence length. We highlight that these exact results need to be treated with caution: the absolute values vary significantly depending on the algorithm hyperparameters chosen, and in the absence of a good external metric we selected these hyperparameters manually via visual inspection of the trajectories and cosine similarities to ensure the results agreed with intuition. However, across different hyperparameters the results consistently showed that non-fixed-point behavior is rare and that the long persona prompt results in a greater rate of non-fixed-point behavior. We leave in-depth classification and explanation of this phenomena to future work.

16

A Mechanistic Analysis of Looped Language Models

Algorithm 1 Per-token limiting behaviour classification Require: similarity (or norm) series s ∈ R𝑛 for one token over the last 𝑛 recurrent firings (final firing excluded); threshold 𝜏; fixed-point fraction 𝜌 Ensure: Label ℓ ∈ {F IXED P OINT, O RBIT, S LIDER, U NKNOWN } 1: — Detrend — ˆ ← arg min𝑎,𝑏 Í𝑖 (𝑠𝑖 − 𝑎𝑖 − 𝑏) 2 2: Fit linear trend: [ 𝑎, ˆ 𝑏] ˆ 3: s̃ ← s − ( 𝑎ˆ t + 𝑏) ⊲ t = [0, 1, . . . , 𝑛 − 1] ⊤ 4: — Spectral amplitude — 5: w ← s̃ ⊙ Hann(𝑛) ⊲ reduce spectral leakage 6: M ← | RFFT(w)| 1: ⊲ discard DC bin 7: 𝑘 ∗ ← arg max 𝑘 𝑀 𝑘 8: 𝐴 ← 4 𝑀 𝑘 ∗ /𝑛 ⊲ Hann-corrected amplitude 9: — Classify (first match wins) — 10: 𝑛close ← #{𝑖 : 𝑠𝑖 ≥ 1 − 𝜏} ⊲ (or 𝑠𝑖 ≤ 𝜏 for norm) 11: if 𝑛close ≥ 𝜌 𝑛 then 12: return F IXED P OINT 13: end if 14: ⊲ peak-to-peak = 2𝐴 ≥ 𝜏; at least 2 full cycles in window 15: if 𝐴 ≥ 𝜏/2 and (𝑘 ∗ + 1)/𝑛 ≥ 2/𝑛 then 16: return O RBIT (freq = (𝑘 ∗ + 1)/𝑛, amp = 𝐴) 17: end if 18: 𝑔 ← 𝑎ˆ ⊲ linear-fit slope; negate for norm series 19: ⊲ sim increases by ≥ 𝜏 over the full window 20: if 𝑔 > 𝜏/𝑛 then 21: return S LIDER (𝑔) 22: end if 23: return U NKNOWN

Table 3. Percentage of tokens exhibiting each behavior type, by model and system prompt.

Model

System Prompt

Huginn Huginn Huginn Huginn Retrofitted-Llama Retrofitted-Llama Retrofitted-Llama Retrofitted-Llama

Long Persona Long Persona (Padded) No System Prompt Short Math Long Persona Long Persona (Padded) No System Prompt Short Math

Non-Fixed-Point %

Orbit %

Slider %

Unknown %

0.14 0.05 0.02 0.01 0.00 0.00 0.00 0.00

0.13 0.02 0.01 0.00 0.00 0.00 0.00 0.00

0.00 0.02 0.01 0.00 0.00 0.00 0.00 0.00

0.01 0.01 0.00 0.00 0.00 0.00 0.00 0.00

Table 4. Percentage of examples exhibiting each behavior type at least once on any question token, by model and system prompt.

Model

System Prompt

Huginn Huginn Huginn Huginn Retrofitted-Llama Retrofitted-Llama Retrofitted-Llama Retrofitted-Llama

Long Persona Long Persona (Padded) No System Prompt Short Math Long Persona Long Persona (Padded) No System Prompt Short Math

Non-Fixed-Point %

Orbit %

Slider %

Unknown %

2.81 0.83 0.76 0.45 0.00 0.00 0.08 0.00

2.50 0.23 0.53 0.38 0.00 0.00 0.08 0.00

0.15 0.61 0.15 0.08 0.00 0.00 0.00 0.00

1.06 0.15 0.08 0.00 0.00 0.00 0.00 0.00

17

A Mechanistic Analysis of Looped Language Models

Ouro (largest amplitude token) Raw signal

FFT magnitudes Orbit (amp=0.166, freq=0.125)

Detrended + windowed (lookback window only)

1.0

detrended * Hann window

0.20

0.14

Hann-corrected amplitude

0.10 Cosine sim to final state

threshold=0.025 dominant freq=0.125

0.16

0.15

0.9

0.8 0.05 0.7 0.00 0.6

-0.05

0.5

-0.10

0

20

40

60

80

120

0.08

0.06

0.02

-0.20 100

0.10

0.04

-0.15

full signal lookback (n=32)

0.4

0.12

0

5

10

15

20

25

0.00

30

0.0

0.2

0.3

Iteration (within lookback)

Frequency (cycles/iter)

Detrended + windowed (lookback window only)

FFT magnitudes Orbit (amp=0.034, freq=0.344) 0.035

0.04

0.03

0.4

0.5

0.4

0.5

threshold=0.025 dominant freq=0.344

1.00 0.030

Hann-corrected amplitude

0.02

0.95 Cosine sim to final state

0.1

Iteration Retrofitted (largest amplitude token) Raw signal

0.01 0.90

0.00

-0.01 0.85 -0.02

0.025

0.020

0.015

0.010

-0.03 0.80

full signal lookback (n=32) 0

20

40

60 Iteration

80

100

120

0.005

detrended * Hann window

-0.04 0

5

10

15

20

Iteration (within lookback)

25

30

0.000

0.0

0.1

0.2

0.3

Frequency (cycles/iter)

Figure 13. Visualizing the component parts of the Orbit detection algorithm of Algorithm 1. The input sequence (cosine similarities for the residual streams of a given token and layer and successive recursions, as compared to their final residual stream) is visualized in the leftmost column. The center column visualizes the effect of windowing and de-trending, and the rightmost column shows the FFT magnitudes. The top row visualizes the detected Orbit with the largest amplitude for the Huginn-0125 model (which occurs with the “Long Persona” prompt) and the bottom row visualizes the largest amplitude for the Retrofitted Llama model (which occurs with no system prompt). We note the appearance of complex, multi-frequency oscillation in this latter case.

18

A Mechanistic Analysis of Looped Language Models

C.2. Do Intermediate Layers Exhibit Non-Fixed-Point Behavior? Geiping et al. (2025) observe orbits and sliders in the latent states after each application of the entire recurrent block. Here we investigate what occurs in the latent states of the intermediate layers within the recurrent block. We first extend Figure 16 of Geiping et al. (2025) (visualizing latent trajectories on a math prompt), visualizing also the trajectories of the intermediate layer residual streams: this can be found in Fig. 15. This demonstrates that in this particular case, where Orbits occur, they also occur in the intermediate layers. This implies – viewed in realized depth – latent trajectories exhibiting multi-scale cyclic behavior. We investigate this across all GSM8k test examples in Fig. 14 by plotting the conditional probability of observing each behavior on a given token for any layer, given an observation of observing another behavior on that same token. We discover that orbits and sliders do not co-occur across looped layers, but that orbits and sliders both frequently co-occur with fixed point behavior. “Unknown” behavior often co-occurs with orbits, which we suggest may be due to mis-classification of the behavior algorithm. Conditional Behavior Co-occurrence: P(Y | X)

FixedPoint

1.00

Orbit

0.00

1.00

0.00

0.26

Slider

0.00

0.00

1.00

0.00

Unknown

0.36

0.36

0.86 0.8

0.6

0.4

Fraction of X where Y also occurs

...Fraction of time Behavior Y also Occurred

1.0

0.2 0.00

0.07

0.00

1.00

FixedPoint

Orbit

Slider

Unknown

0.0

Given Behavior X Occurred...

Figure 14. Conditional probabilities of co-occurrence for the different limiting behaviors.

19

A Mechanistic Analysis of Looped Language Models Token: "Cla" L2

L3

L4

L5

6

6

6

5

0

0

0

0

-6 -5

0

5

-6 -5

Token: "ire" L2

0

5

-6 -5

L3

0

5

-5 -5

L4

5

4

3

4

0

0

0

0

-5 -7

0

7

-4 -9

Token: " makes" L2

0

9

-3 -9

L3

0

9

-4 -8

L4

4

4

4

0

0

0

0

0

17

-4 -17

Token: " a" L2

0

17

-4 -18

L3

0

18

-4 -15

L4

3

3

3

0

0

0

0

0

17

-3 -17

Token: " 3" L2

0

17

-3 -17

L3

0

17

-3 -15

L4

2

2

1

0

0

0

0

0

10

-2 -12

0

12

-2 -11

0

8

0

15

0

15

L5

1

-1 -10

0

L5

4

-4 -17

5

L5

3

-3 -17

0

L5

11

-1 -9

0

9

Figure 15. PCA trajectories in the intermediate layers of Huginn-0125: this reproduces the leftmost column of Fig. 16 in Geiping et al. (2025) (the first two principal components) and additionally plots the latent trajectories for the intermediate layers in the recurrent block.

20

A Mechanistic Analysis of Looped Language Models

C.3. How Does Non-Fixed-Point Behavior Impact Stages of Inference? To complete the link between non-fixed-point behavior and our work we additionally investigate how this behavior impacts the observed stages of inference. To attempt to isolate the “worst case” scenario for stages of inference stability, we plot stages of inference for the GSM8k test prompt that exhibits the greatest orbit amplitude. We first plot in realized depth the extended stages of inference metrics used throughout the paper, see Fig. 16. This demonstrates that sink rates show some variability with the orbiting behavior, but the other metrics remain broadly consistent. We plot also the same metrics in block depth (showing only the final 64 loops) in Fig. 17: this demonstrates clearly that the stages of inference remain very consistent despite the orbiting behavior.

0.4 0.2

0.010

0.008

0.006

0.4

Layer

Mixing Score

Sink Rate

0.6

Colsum Concentration

Late 0.5

0.3 0.2

0.0 0

100

200

300

400

500

0

100

Recurrent Position

200

300

400

500

0

Recurrent Position

100

200

300

400

500

Early

Recurrent Position

Figure 16. Stages of inference metrics (sink rate, mixing score and colsum concentration) for Huginn-0125 across 128 recurrences, for the GSM8k test prompt that exhibited the largest orbit amplitude. Visualized in realized depth.

0.4

0.2

0.012 0.010 0.008

0

20

40

60

% Block Depth

80

100

0

20

40

60

% Block Depth

80

100

0.40 Recurrence

Mixing Score

0.6 Sink Rate

Colsum Concentration

Late 0.014

0.35 0.30 0.25 0.20 0

20

40

60

80

100

Early

% Block Depth

Figure 17. Stages of inference metrics (sink rate, mixing score and colsum concentration) for Huginn-0125 across 128 recurrences, for the GSM8k test prompt that exhibited the largest orbit amplitude: only the final 64 recurrences are visualized, to isolate the impact of the orbit. Visualized in percentage block depth.

We plot the same for the largest amplitude orbit on the Retrofitted Llama model in Fig. 18, demonstrating that the same behavior holds.

21

A Mechanistic Analysis of Looped Language Models Late

0.6 0.4 0.2

0.020

0.015

0.010

0.0 0

20

40

60

80

100

0

% Block Depth

20

40

60

80

0.6 Recurrence

Mixing Score

Sink Rate

0.8

Colsum Concentration

1.0

0.4

0.2

100

0

20

% Block Depth

40

60

80

100

Early

% Block Depth

Figure 18. Stages of inference metrics (sink rate, mixing score and colsum concentration) for Retrofitted Llama across 128 recurrences, for the GSM8k test prompt that exhibited the largest orbit amplitude: only the final 64 recurrences are visualized, to isolate the impact of the orbit. Visualized in percentage block depth.

D. Additional Fixed Point Results D.1. Cyclic Similarity We include here additional plots to validate our cyclic similarity claims in Sec. 4. Fig. 19 complements Fig. 2 by visualizing the cosine similarity between the residual streams after each layer for the range of Looped Transformers visualized in the original figure. Additional models (Huginn-0125 and all retrofitted models) are visualized in Fig. 19 for 32 recurrences, demonstrating that the cyclic similarity is consistent for larger numbers of recurrences. We note that Huginn-0125 continues its previous trend of all layer outputs converging to similar representations, with some cyclic similarity still visible. Fig. 21 similarly provides an extended version of Fig. 2, demonstrating that attention matrix cyclic similarity holds to 32 recurrences. Retrofitted Llama

1.0

0

Huginn-0125

1.0

0

6

72

0.6

96 0.4

120 144

0.2

168 0

24 48 72 96 120 144 168 Layer Index

0.0

Cosine Similarity

0.8

48

4

0.8

12 18

0.6

24 30

0.4

36 42

0.2

48

Cosine Similarity

24

1.0

0.8

8 12

0.6

16 20

0.4

24 28

Cosine Similarity

Ouro 1.4B 0

0.2

32

54 0

6 12 18 24 30 36 42 48 54 Layer Index

0.0

0

4

8

12 16 20 24 28 32

0.0

Layer Index

Figure 19. Cosine similarity between residual streams after every pair of layers for different Transformer models, averaged across the batch and sequence dimensions. Left: Ouro 1.4B (Zhu et al., 2025). Center: Retrofitted Llama (McLeish et al., 2025). Right: Huginn-0125 (Geiping et al., 2025). All models looped 8 times. Diagonal patterns indicate that the residual stream after each sub-block is most similar to the same block in the next recurrence; every layer in the recurrent block reaches a different fixed point.

22

A Mechanistic Analysis of Looped Language Models

0.2 128 64

128

0.6 0.4

128

0.2 192

0.0

0

Layer Index

64

128

192

0.8 64

0.6 0.4

128

0.2 192

0.0

0

64

Layer Index

128

192

0

0.8

64

0.6

128

0.4 192

0.2

256

0.0

1.0

0

Layer Index

64

128 192 256

Cosine Similarity

0.4

0.8 64

Retrofitted TinyLlama

1.0

0

Cosine Similarity

0.6

64

Retrofitted OLMo

1.0

0 Cosine Similarity

0.8

0

Retrofitted Llama

1.0

Cosine Similarity

Huginn-0125 0

0.0

Layer Index

Figure 20. Cosine similarity between residual streams after every pair of layers for different Transformer models, averaged across the batch and sequence dimensions. Left: Huginn-0125 (Geiping et al., 2025). Center Left: Retrofitted Llama. Center Right: Retrofitted OLMo. Right: Retrofitted TinyLlama (McLeish et al., 2025). All models looped 32 times. Extended version of Fig. 19.

4 2 128 0

64

6

64

4 128 2 192

128

0

Layer Index

64

128

Retrofitted TinyLlama 10

0

8 64

6 4

128

2 192

192

0

64

Layer Index

128

192

0

10

64

8 6

128

4

192

2 256 0

Layer Index

64

Avg. Frobenius Norm

6

64

Avg. Frobenius Norm

8

Retrofitted OLMo

8

0

Avg. Frobenius Norm

Retrofitted Llama 10

Avg. Frobenius Norm

Huginn-0125 0

128 192 256

Layer Index

Figure 21. Frobenius norm between attention matrices for different Transformer models, averaged across the batch and head dimensions. Left: Huginn-0125 (Geiping et al., 2025). Center Left: Retrofitted Llama. Center Right: Retrofitted OLMo. Right: Retrofitted TinyLlama (McLeish et al., 2025). All models looped 32 times. Extended version of Fig. 2.

D.2. Fixed Point and Successive Differences

Ouro 1.4B (128 steps)

Retrofitted Llama (128 steps)

1.0

Huginn-0125 (128 steps)

1.0

0.8

0.8

0.9

0.6

Late

1.0

Layer

Fixed Point Cosine Similarity

In Sec. 4 we demonstrate that retrofitted Llama and Huginn-0125 reach a fixed point but Ouro does not. We do so by – for each layer – computing an “approximate fixed point” after 128 recurrences and then plotting the norm of the difference for each layer at successive recurrences to this fixed point. In Fig. 22 we demonstrate that this same behavior is observed when considering cosine similarity to the fixed point, and in Fig. 23 we see that it holds for attention matrices.

0.6

0.4

0.8

0.4

0.2 0

20

40

60

80

Recurrence

100

120

0

20

40

60

80

Recurrence

100

120

0

20

40

60

80

100

120

Early

Recurrence

Figure 22. Cosine similarity between the residual stream after each layer in the recurrent block and its “approximate fixed point” the residual stream after that layer in the 128th recursion. While Huginn-0125 and retrofitted Llama quickly reach a fixed point, Ouro does not - even though the cosine similarity between successive recursions tends towards one, as evidenced by Fig. 22.

We then demonstrated in Fig. 3 that all models, independent of whether they reach a strict fixed point or not, demonstrate converge towards very low difference norms between successive residual streams. We verify this same behavior holds for residual stream cosine similarities and attention matrix Frobenius norms in Figs. 24 and 25 respectively.

23

Ouro 1.4B (128 steps)

Retrofitted Llama (128 steps)

Huginn-0125 (128 steps) Late

15 10 5 0

4

8

3

6

2

4

1

2

0 0

20

40

60

80

100

120

Layer

Fixed Point Attn. Diff. Norm.

A Mechanistic Analysis of Looped Language Models

0 0

20

40

Recurrence

60

80

100

120

0

20

40

Recurrence

60

80

100

120

Early

Recurrence

Ouro 1.4B (128 steps)

Retrofitted Llama (128 steps)

Huginn-0125 (128 steps)

1.0

1.00

1.0

0.8

0.95

0.9

0.90

0.8

0.6

0.85

0.4

0.7

0.80

0.6

0.2 0

20

40

60

80

100

Late

Layer

Successive Cosine Similarity

Figure 23. Frobenius norm between attention matrices of each layer in the recurrent block and their corresponding “approximate fixed point”; the attention matrices of the same layer in the 128th recursion. While Huginn-0125 and retrofitted Llama quickly reach a fixed point, Ouro does not.

120

0

20

40

Recurrence

60

80

100

120

0

20

40

Recurrence

60

80

100

120

Early

Recurrence

Ouro 1.4B (128 steps)

Retrofitted Llama (128 steps)

Huginn-0125 (128 steps) Late

15

8

4 10 5 0

3

6

2

4

1

2 0

0 0

20

40

60

80

Recurrence

100

120

Layer

Successive Attn. Diff. Norms

Figure 24. Cosine similarities between successive recursions of the residual stream after the same layer.

0

20

40

60

80

Recurrence

100

120

0

20

40

60

80

100

Recurrence

Figure 25. Frobenius norm between attention matrices of each layer between successive recurrences.

24

120

Early

A Mechanistic Analysis of Looped Language Models

D.3. Latent Space Trajectories In this section we visualize additional latent space trajectories, supplementing Fig. 5. All trajectories are plotted by taking all latent states of the final token position on the test sequence described in App. B, computing PCA on these latent state vectors and resultant dimensional reduction of this sequence of vectors. They are intended as illustrations to demonstrate qualitative behavior. Ouro is visualized in Fig. 26 and Huginn-0125 in Fig. 27. The Ouro trajectory is of particular interest as we are aware from the results in the main body of the paper that this model does not reach a strict fixed point. Fig. 26 demonstrates that – on this test sequence – Ouro reaches an approximately constant trajectory, but with visibly larger deviations even at later recurrences than Huginn-0125 in Fig. 27. However, we find evidence that this “approximately stable trajectory” behavior is not universal: for a simple “maths” test prompt (The square root of 16 is) shown in Fig. 28, Ouro appears to reach a stable trajectory in recurrences 8-16 before departing from this and appearing to become “unstable”. Recurrences 0-8

Recurrences 8-16

Recurrences 16-24

Recurrences 24-32

5

5

5

0

0

0

0

-5

-5

-5

-5

-10

-10 0

10

-10

20

0

10

PC 1

Recurrence

PC 2

Late 5

-10

20

0

10

PC 1

20

0

10

PC 1

20

Early

PC 1

Figure 26. Ouro 1.4B (Zhu et al., 2025) latent space trajectory traced out by the hidden states of the final sequence position on the test prompt; reduced to two dimensions by computing PCA over all final sequence position embeddings. Recurrences 0-8

Recurrences 8-16

Recurrences 16-24

Recurrences 24-32

0

0

0

0

-20

-20

-20

-20

-40

-40

-40

-40

-60

-60 -40

-20

0

20

-60 -40

-20

PC 1

0

20

Recurrence

PC 2

Late

-60 -40

-20

PC 1

0

20

-40

-20

PC 1

0

20

Early

PC 1

Figure 27. Huginn-0125 (Geiping et al., 2025) latent space trajectory traced out by the hidden states of the final sequence position on the test prompt; reduced to two dimensions by computing PCA over all final sequence position embeddings. Recurrences 8-16

Recurrences 16-24

Recurrences 24-32

15

15

15

10

10

10

10

5

5

5

5

0

0

0

0

-5

0

10 PC 1

20

-5

0

10

20

-5

PC 1

0

10 PC 1

20

-5

Late

Recurrence

PC 2

Recurrences 0-8 15

0

10

20

Early

PC 1

Figure 28. Ouro 1.4B (Zhu et al., 2025) latent space trajectory traced out by the hidden states of the final sequence position on a “maths” test prompt (The square root of 16 is); reduced to two dimensions by computing PCA over all final sequence position embeddings.

25

A Mechanistic Analysis of Looped Language Models

D.4. Architecture Choices Here we provide more complete results to supplement Sec. 4.2, visualizing fixed point difference norms and cosine similarities for various architecture choices in Figs. 29 and 30 respectively.

Fixed Point Norm

Pre Norm Inp. Inj.

Huginn-0125 Norm Inp. Inj.

20

20

10

10

0

100

50

0

0 0

10

20

30

40

50

60

0

10

20

Recurrence Pre Norm Fixed Point Norm

Ouro Norm Inp. Inj.

1500

Distance to self Distance to Max

1000

30

40

50

60

0

10

20

30

40

Recurrence

Recurrence

Huginn-0125 Norm

Ouro Norm

50

60

50

60

30 100 20

500

50

10 0

0 0

10

20

30

40

50

60

0 0

10

20

Recurrence

30

40

50

60

0

10

20

Recurrence

30

40

Recurrence

Figure 29. Norm of the difference between the residual stream after the first layer in the recurrent block for successive recurrences and an “approximate fixed point” – the residual stream in the 128th recurrence. Two fixed point differences are visualized: the difference to the fixed point of the same (first) layer (blue) and the difference to the fixed point which has the greatest norm difference from the first layer (green). Visualized are a range of norm structures (columns), with input injection (top row) and without (bottom row). All models are randomly initialized with 12 layers in the recurrent block, and no prelude or coda.

Fixed Point Cosine Sim

Pre Norm Inp. Inj.

Huginn-0125 Norm Inp. Inj. 1.0

1.00

0.8

0.8

0.75

0.6

0.6

0.4

0.50 0.25

0.4

0.00 0

10

20

30

40

50

60

0

10

20

Recurrence

30

40

50

60

0

Huginn-0125 Norm 1.00

0.75

0.75

0.75

0.50

0.50

0.50

0.25

0.25

Similarity to self Similarity to Min 0

10

20

30

40

Recurrence

50

60

30

40

50

60

50

60

Ouro Norm

1.00

0.00

20

Recurrence

1.00

0.25

10

Recurrence

Pre Norm Fixed Point Cosine Sim

Ouro Norm Inp. Inj.

1.0

0.00

0.00 0

10

20

30

40

Recurrence

50

60

0

10

20

30

40

Recurrence

Figure 30. Cosine similarity between the residual stream after the first layer in the recurrent block for successive recurrences and an “approximate fixed point” – the residual stream in the 128th recurrence. Two fixed point differences are visualized: the difference to the fixed point of the same (first) layer (blue) and the difference to the fixed point which has the lowest cosine similarity to the first layer (green). Visualized are a range of norm structures (columns), with input injection (top row) and without (bottom row). All models are randomly initialized with 12 layers in the recurrent block, and no prelude or coda.

We also verify that the results presented are not particular to 12 layers by visualising results for both 4 and 16 layers in Fig. 31, showing qualitatively identical behavior.

26

A Mechanistic Analysis of Looped Language Models Fixed Point Cosine Sim

Pre Norm Inp. Inj.

Huginn-0125 Norm Inp. Inj. 1.0

1.00

0.8

0.8

0.75

0.6

0.6

0.50

0.4

0.4

0.2

0.25 0.00

0

10

20

30

40

50

60

0

10

20

Recurrence

30

40

50

60

0

10

20

Recurrence

Pre Norm Fixed Point Cosine Sim

Ouro Norm Inp. Inj.

1.0

Huginn-0125 Norm

1.00

30

40

50

60

50

60

Recurrence Ouro Norm 1.0

1.0

0.75 0.5

Similarity to self Similarity to Min 4 layers 16 layers

0.50 0.25 0.00 0

10

20

30

40

50

0.5

0.0

60

0.0 0

10

Recurrence

20

30

40

50

60

0

10

20

Recurrence

30

40

Recurrence

Figure 31. Cosine similarity between the residual stream after the first layer in the recurrent block for successive recurrences and an “approximate fixed point” – the residual stream in the 128th recurrence. Two fixed point differences are visualized: the difference to the fixed point of the same (first) layer (blue) and the difference to the fixed point which has the lowest cosine similarity to the first layer (green). Visualized are a range of norm structures (columns), with input injection (top row) and without (bottom row). Here models of 4 and 16 layers are compared, showing qualitatively identical behavior and demonstrating that the behavior is not particular to 12 layers.

E. Additional Stages of Inference Results In this appendix we explore in more detail the mixing behavior presented in the main body via additional metrics, and considering a wider range of models. E.1. Input Independent Metrics As our work is concerned largely with how the functionality of layers change throughout depth as the residual stream is iteratively updated, our primary concern is with input-dependent measures of stages of inference: ColSum concentration as presented in the main text of the paper is one such input-dependent metric, and the later sections of this appendix will introduce and present results for a wider range of such metrics. However, to bridge the gap between our work and that of Lad et al. (2024), we also present results for the fraction of prediction and suppression neurons (Gurnee et al., 2024) in successive layers of both feedforward and looped Transformers. We highlight however that these are unable to change with successive recurrences, and are thus secondary to our focus. Fig. 32 presents results for a selection of feedforward models used throughout the paper, and Fig. 33 presents results for a selection of looped models used throughout the paper. Similar to our results elsewhere, we find that the looped model stages of inference in these metrics tend to mirror those of feedforward models. GPT-2

Llama 3.2 1B

TinyLlama 1.1B

OLMo-2 1B

Fraction of neurons

0.020 0.015

0.15 0.100

0.08 0.10

0.06

0.010

0.075 0.050

0.04 0.05

0.005 0.000

Prediction Suppression

0.125

0.10

0.025

0.02 0

25

50 Depth (%)

75

100

0.00

0

25

50

75

100

Depth (%)

0.00

0

25

50 Depth (%)

75

100

0.000

0

25

50

75

Depth (%)

Figure 32. Fraction of prediction and suppression neurons in a selection of feedforward models used throughout the paper.

27

100

A Mechanistic Analysis of Looped Language Models Ouro 1.4B 0.04 pre Fraction of neurons

Retrofitted Llama

Huginn-0125

0.15

core

coda

0.10 pre

core

Retrofitted OLMo

coda

coda

0.15

pre

0.06

0.10

0.10

0.05

0.05

core coda Prediction Suppression

0.02 0.04

0.05 0.00

core

0.08

0.03

0.10

0.15 pre

Retrofitted TinyLlama

0.01

0

50

100

0.00

0.02 0

Depth (%)

50

100

0.00

0

50

Depth (%)

100

0.00

0

Depth (%)

50

100

0.00

Depth (%)

0

50

100

Depth (%)

Figure 33. Fraction of prediction and suppression neurons in a selection of looped models used throughout the paper.

E.2. Input Dependent Metrics One well-studied phenomenon by which Transformers drastically reduce the mixing in given layer is that of the attention sink (Xiao et al., 2023; Barbero et al., 2025), whereby the layer focuses the majority of the attention “weight” onto the first token in the sequence; often this is the BOS (beginning of sequence) token, and thus completely uninformative. To measure this behavior, we adopt the attention sink score and sink rate of Gu et al. (2024): for token position 𝑘 and sequence length 𝑇, the sink score at layer ℓ and head ℎ is defined as: 𝑇 −1

sink-score 𝑘(ℓ,ℎ) =

1 ∑︁ (ℓ,ℎ) 𝐴 , 𝑇 𝑡=0 𝑡 𝑘

(16)

where 𝐴𝑡(ℓ,ℎ) corresponds to the realized attention matrix of Eq. (2) at layer ℓ, head ℎ. The sink rate is then defined as the 𝑘 fraction of heads for which the sink score lies above a certain threshold: 1 sink-rate 𝑘(ℓ ) = 𝐻

𝐻   ∑︁ I sink-score 𝑘(ℓ,ℎ) ≥ 𝜏 ,

(17)

ℎ=1

where following related work we define a threshold of 𝜏 = 0.3, and I denotes the indicator function. We also adopt the Mixing score of Queipo-de Llano et al. (2025): for sequence length 𝑇, layer ℓ and head ℎ this is defined as the average row entropy of the attention matrices: 𝑇

mixing-score (ℓ,ℎ) =

1 ∑︁ (ℓ,ℎ) 𝐻 ( 𝐴𝑖,: ). 𝑇 𝑖=1

(18)

Following Skean et al. (2025); Queipo-de Llano et al. (2025) we additionally measure the compression of the residual stream X via the matrix-based entropy 𝐻 (X). In the plots that follow we average sink rates, Mixing scores and ColSum concentrations over all the heads in a layer, and over all input sequences. Residual entropy is averaged over all input sequences. E.3. Cyclic Stages of Inference Supplementing the cyclic recurrence in realized depth for retrofitted Llama in Fig. 7, we additionally overlay cycles in realized depth for Ouro and Huginn-0125 in Fig. 34, where all models are run for 8 recurrences. This demonstrates that for sink rates and mixing scores, the cyclic behavior and lack of significant depth-wise changes hold both for the other investigated models, and the additional stages of inference metrics. The results for residual entropy are more nuanced: while the results for Huginn-0125 and the retrofitted models show broadly the same behavior as the attention-based metrics, Ouro demonstrates different behavior. As shown in Fig. 35, the residual entropies both change significantly with successive recurrences and do not closely follow the feedforward behavior. We believe that the divergence from feedforward behavior may derive from the norm structure of this model: the residual stream 28

A Mechanistic Analysis of Looped Language Models

is normalised after each recurrent block, as visualised in Fig. 9, thus periodically shutting down massive activations. This is not the case for the retrofitted series of models, which lack this norm and show much closer alignment with feedforward stages of inference. Ouro 1.4B

Huginn-0125

Llama 3.2 1B

0.8 Colsum Concentration

0.0200 0.0175 Mixing Score

0.8 0.6 0.4 0.2

0.0150 0.0125 0.0100 0.0075 0.0050

0

50

100

0.6 0.5 0.4 0.3

3

2

1

0.2

0.0025

0.0

4

0.7 Residual Entropy

1.0

Sink Rate

Retrofitted Llama

0 0

% Recurrent Position

50

100

0

% Recurrent Position

50

100

0

% Recurrent Position

50

100

% Recurrent Position

Figure 34. Stages of inference for a selection of Looped transformers, all using 8 recurrences: Huginn-0125 (Geiping et al., 2025), Ouro 1.4B (Zhu et al., 2025) and Llama with retrofitted recurrences (McLeish et al., 2025). Note Huginn-0125 and Retrofitted Llama have prelude and coda layers too: each 2 layers in Huginn-0125 and each 4 in Retrofitted Llama.

For completeness, we plot these stages of inference for all other models referenced in the paper. See Fig. 35 (Ouro 1.4B), Fig. 36 (Huginn-0125), Figs. 37 to 39 (retrofitted Llama, OLMo and TinyLlama). We plot stages of inference for Ouro 2.6B in Fig. 40. This model is interesting due to the training regime followed by Zhu et al. (2025), which “upcycles” a 48 layer model from the 24 layer 1.4B parameter model. As a consequence, the first and second half of each recurrent block each independently align with the Llama feedforward stages of inference. In Sec. 5 we suggested that the lack of stages of inference in Huginn-0125 is likely due to the normalization of the residual stream resulting in massive activations being unable to form. Here we further support this suggestion by ablating the massive activations from the Retrofitted Llama model (which does display stages of inference) via zeroing the output of the MLP in the second layer, which is responsible for its massive activations. In this setting, visualized in Fig. 41, we see that the model no longer exhibits stages of inference comparable to the feedforward model, suggesting that the presence of massive activations is required for stages of inference to emerge in looped models.

R2

0.6 0.4 0.2

Colsum Concentration

Mixing Score

0.8 Sink Rate

R1

0.020 0.015 0.010 0.005

0.0 0

50 % Block Depth

100

0

50

100

R3 0.8

4

0.6 0.4 0.2

3 2 1 0

0

% Block Depth

Llama 3.2 1B Residual Entropy

R0 1.0

50 % Block Depth

100

0

50

100

% Block Depth

Figure 35. Stages of inference for each recurrent loop in Ouro 1.4B. The close overlap with feedforward stages of inference is a particularly striking result as this model is trained from scratch with recurrence.

29

A Mechanistic Analysis of Looped Language Models

Prelude Mixing Scores

Mixing Score

0.6 0.4

Colsum Concentration

0.020

0.8 Sink Rate

Colsum Concentrations

0.015

0.010

0.2 0.005 50

100

0

% Block Depth

50

Residual Entropy Late

0.7 0.6 0.5 0.4 0.3 0.2

4 Recurrence

1.0

0

Llama 3.2 1B

Residual Entropy

Sink Rates

Coda

3 2 1 0

100

0

% Block Depth

50

100

0

% Block Depth

50

100

Early

% Block Depth

Figure 36. Stages of inference for each recurrent loop in Huginn-0125. This represents a negative result: stages of inference do not occur. We discuss possible causes for this in the main text.

Prelude Sink Rates

Llama 3.2 1B

Mixing Scores

Colsum Concentrations

Residual Entropy Late

0.6

0.015

0.010

0.4 0.005 0

50

100

0.7

2.5

0.6 0.5 0.4 0.3

% Block Depth

50

100

1.5 1.0 0.5 0.0

0.2 0

2.0

Recurrence

0.8

Colsum Concentration

Mixing Score

0.020

Residual Entropy

1.0

Sink Rate

Coda

0

% Block Depth

50

100

0

% Block Depth

50

100

Early

% Block Depth

Figure 37. Stages of inference for each recurrent loop in the retrofitted Llama model (McLeish et al., 2025). Each block demonstrates very similar stages of inference to Llama, the base model from which pretrained layers are taken.

Prelude Sink Rates

Coda

Mixing Scores

Colsum Concentrations

Residual Entropy Late

0.6

0.4 0.2

0.015

0.010

0.0 0

50 % Block Depth

100

0

50

0.5 0.4 0.3 0.2

4.0 Recurrence

0.6

0.020

Residual Entropy

Mixing Score

0.8

Colsum Concentration

1.0

Sink Rate

OLMo2

3.5

3.0

0.1

100

0

% Block Depth

50 % Block Depth

100

0

50

100

Early

% Block Depth

Figure 38. Stages of inference for each recurrent loop in the retrofitted OLMo model (McLeish et al., 2025). Similarly, each block demonstrates very similar stages of inference to OLMo, the base model from which pretrained layers are taken.

30

A Mechanistic Analysis of Looped Language Models

Prelude Mixing Scores

Colsum Concentrations

0.6 0.4 0.2

0.020 0.015 0.010 0.005

Residual Entropy Late

0.8 0.6 0.4 0.2

0.0

3 Recurrence

0.025 Mixing Score

Sink Rate

Colsum Concentration

1.0 0.8

TinyLlama

Residual Entropy

Sink Rates

Coda

2

1

0 0

50

100

0

50

% Block Depth

100

0

% Block Depth

50

100

0

% Block Depth

50

100

Early

% Block Depth

Figure 39. Stages of inference for each recurrent loop in the retrofitted TinyLlama model (McLeish et al., 2025). Similarly, each block demonstrates very similar stages of inference to TinyLlama, the base model from which pretrained layers are taken.

R0, 1 R0, 2 1.0

R2, 1 R2, 2

R3, 1 R3, 2

Llama 0.9

0.0200

0.8

Mixing Score

0.6 0.4 0.2

Colsum Concentration

0.0175

0.8 Sink Rate

R1, 1 R1, 2

0.0150 0.0125 0.0100 0.0075 0.0050 0.0025

0.0 0

20

40

60

80

100

0.7 0.6 0.5 0.4 0.3 0.2

0

% Recurrent Depth

20

40

60

80

100

0

20

% Recurrent Depth

40

60

80

100

% Recurrent Depth

Figure 40. Stages of inference for each recurrent loop in Ouro 2.6B. For this model we separate out the first and second half of the recurrent block and overlay them, demonstrating that both halves have close alignment with the Llama feedforward stages of inference. We suggest that this likely arises due to the training regime of Zhu et al. (2025), which first trains a single 24 layer 1.4B parameter model, and then “upcycles” this into a 48 layer 2.6B parameter model by duplicating these layers, consequently duplicating the stages of inference as well.

0.25

0.015

0.010

0.00 0.005 0

50 % Block Depth

100

0

50

100

Llama 3.2 1B

Late

0.6 0.4 0.2 0.0

% Block Depth

0

50 % Block Depth

100

2 Recurrence

0.50

Colsum Concentration

Mixing Score

0.75 Sink Rate

Coda

0.020

Residual Entropy

Prelude 1.00

1

0

0

50

100

Early

% Block Depth

Figure 41. Stages of inference for each recurrent loop in Retrofitted Llama for which the massive activations have been ablated.

31

A Mechanistic Analysis of Looped Language Models

E.4. Non-Reasoning Stages of Inference Throughout the rest of the paper, experiments are conducted on the GSM8k dataset. In this appendix we verify that the stages of inference we observe are not specific to this dataset, and also occur in a non-reasoning setting, for which we use the HellaSwag dataset (Zellers et al., 2019). We follow an identical experimental setup to the GSM8k experiments, running inference on 256 random examples from the test split. We present results in Fig. 42 (Ouro 1.4B), Fig. 43 (Huginn-0125), Figs. 44 to 46 (retrofitted Llama, OLMo and TinyLlama). These show very few deviations from the GSM8k results, and the conclusions throughout the rest of the paper hold. However, we highlight the following small differences: • Across the board, sink rates tend to be higher in the HellaSwag setting. • In Retrofitted OLMo-2, ColSum concentration appears to be slightly higher in the HellaSwag setting. Llama 3.2 1B Sink Rates

Mixing Scores

Colsum Concentrations

0.2 0.0

0.03 0.02 0.01

0

50

100

0

% Block Depth

50

0.6

0.4

0.2

100

2.5 Recurrence

0.4

0.04

Residual Entropy

Mixing Score

Sink Rate

0.6

3.0

0.8

Colsum Concentration

0.05

0.8

-0.2

Residual Entropy Late

1.0

2.0 1.5 1.0 0.5 0.0

0

% Block Depth

50

100

0

% Block Depth

50

100

Early

% Block Depth

Figure 42. Stages of inference for each recurrent loop in Ouro 1.4B, run on the HellaSwag dataset. Prelude

0.6 0.4

0.03 0.02

0.2

% Block Depth

100

0.6 0.5 0.4 0.3 0.2

0.01 50

Late

0.7

0

50

100

2 1 0

0

% Block Depth

3 Recurrence

0.04

Residual Entropy 4 Residual Entropy

0.8

Colsum Concentrations Colsum Concentration

0.05

0

Llama 3.2 1B

Mixing Scores

1.0 Mixing Score

Sink Rate

Sink Rates

Coda

50 % Block Depth

100

0

50 % Block Depth

Figure 43. Stages of inference for each recurrent loop in Huginn-0125, run on the HellaSwag dataset.

32

100

Early

A Mechanistic Analysis of Looped Language Models

Prelude Sink Rates

Llama 3.2 1B

Mixing Scores

Colsum Concentrations

Residual Entropy Late

2.0

0.8 0.7 0.6 0.5

0.04 0.03 0.02

0.7 0.6 0.5 0.4 0.3

0.01 0

50

100

0

% Block Depth

50

100

1.5 Recurrence

Mixing Score

0.9

Colsum Concentration

0.05

Residual Entropy

1.0

Sink Rate

Coda

1.0 0.5 0.0

0

% Block Depth

50

100

0

% Block Depth

50

100

Early

% Block Depth

Figure 44. Stages of inference for each recurrent loop in the retrofitted Llama model (McLeish et al., 2025), run on the HellaSwag dataset.

Prelude Sink Rates

Mixing Scores

Colsum Concentrations

0.2

0.05 0.04 0.03 0.02

0.0

Late 3.5

0.6 0.5 0.4 0.3 0.2

Recurrence

0.4

Residual Entropy

Residual Entropy

0.6

Colsum Concentration

Mixing Score

0.8

OLMo2

0.7

0.06

1.0

Sink Rate

Coda

3.0 2.5 2.0

0.1 0

50

100

0

% Block Depth

50

100

0

% Block Depth

50

100

0

% Block Depth

50

100

Early

% Block Depth

Figure 45. Stages of inference for each recurrent loop in the retrofitted OLMo model (McLeish et al., 2025), run on the HellaSwag dataset. ColSum concentration deviates slightly from its GSM8k counterpart here, but still broadly follows the same stages of inference as the feedforward OLMo model.

Prelude Sink Rates

Coda

Mixing Scores

Colsum Concentrations

Residual Entropy Late

0.4 0.2

0.06

0.04

0.02

0.0

3

0.6 0.4 0.2

Recurrence

0.6

0.8 Residual Entropy

Mixing Score

0.8

Colsum Concentration

1.0

Sink Rate

TinyLlama

2

1

0 0

50 % Block Depth

100

0

50

100

0

% Block Depth

50 % Block Depth

100

0

50

100

Early

% Block Depth

Figure 46. Stages of inference for each recurrent loop in the retrofitted TinyLlama model (McLeish et al., 2025), run on the HellaSwag dataset.

33

A Mechanistic Analysis of Looped Language Models

E.5. Stability To Unseen Test-Time Recurrences This section extends the results presented in Sec. 5.2. We supplement Fig. 11 by plotting how stages of inference change per-layer throughout recurrences for additional models: these can be found in Figs. 47 to 49. The large standard deviations in Huginn-0125 and retrofitted Llama mixing scores reflect the fact that these models tend to reach different, but still stable, constant states.

0.50 0.25

0.015 0.010 0.005

0.00 0

1000

2000

3000

0

Recurrent Position

1000

2000

0.6 0.4 0.2

3000

Late

4

0

Recurrent Position

1000

2000

3 Layer

Mixing Score

Sink Rate

0.75

0.8 Residual Entropy

Colsum Concentration

0.020 1.00

2 1

3000

0

Recurrent Position

1000

2000

3000

Early

Recurrent Position

Figure 47. Stages of inference for each of the distinct blocks in Ouro (Zhu et al., 2025), as they are reapplied throughout the model for 128 recurrences. These consistently change throughout the realized depth of the model, reaching no clear fixed point. Mean and standard deviation are over separate inputs to the model, taken over the GSM8k subset. Late

0.2

0.0125 0.0100 0.0075 0.0050

0

200

0.5 0.4 0.3

0

Recurrent Position

200

4.2 4.0 3.8 3.6

0.2

400

4.4 Layer

0.4

0.0150

Residual Entropy

Mixing Score

Sink Rate

0.6

Colsum Concentration

0.0175

400

0

Recurrent Position

200

400

0

Recurrent Position

200

Early

400

Recurrent Position

Figure 48. Stages of inference for each of the distinct blocks in Huginn-0125 (Geiping et al., 2025), as they are reapplied throughout the model for 128 recurrences. These converge to constant behavior. Mean and standard deviation are over separate inputs to the model, taken over the GSM8k subset.

0.7 0.6 0

250

500

Recurrent Position

750

0.015

0.010

0

250

500

0.5

0.4

0.3

750

0

Recurrent Position

250

500

Recurrent Position

750

0.15 Layer

Mixing Score

Sink Rate

0.8

Residual Entropy

0.020

0.9

0.5

Colsum Concentration

Late

1.0

0.10

0.05 0

250

500

750

Early

Recurrent Position

Figure 49. Stages of inference for each of the distinct blocks in retrofitted Llama (McLeish et al., 2025), as they are reapplied throughout the model for 128 recurrences. These converge to constant behavior. Mean and standard deviation are over separate inputs to the model, taken over the GSM8k subset.

We additionally plot the extended versions of Fig. 12 in Figs. 50 to 52.

34

A Mechanistic Analysis of Looped Language Models Llama 3.2 1B

Sink Rate

0.8

0.020

0.8

0.015

0.6

0.010

0.4

0.005

0.2

Late 3

0.6

Recurrence

1.0

2

0.4

1

0.2 0.0

0 0

25

50

75

100

0

25

% Block Depth

50

75

100

0

25

% Block Depth

50

75

100

0

25

% Block Depth

50

75

100

Early

% Block Depth

Figure 50. Stages of inference for Ouro with each of 128 recurrences, visualized over percentage recurrent depth. These consistently change with successive recurrences, deviating significantly from the stages of inference seen with train-time recurrences.

0.010

0.2 0.005 0

50

100

0

% Block Depth

50

Late 4

0.6 0.5 0.4 0.3 0.2

100

Recurrence

0.4

0.015

Llama 3.2 1B

0.7 Residual Entropy

0.6

Coda Colsum Concentration

Mixing Score

0.8 Sink Rate

Prelude

0.020

1.0

3 2 1 0

0

50

% Block Depth

100

0

50

% Block Depth

100

Early

% Block Depth

Figure 51. Stages of inference for Huginn-0125 with each of 128 recurrences, visualized over percentage recurrent depth. These quickly reach a fixed point and do not deviate far from their starting stages of inference.

0.4

0.010

0.005 0

50

100

0

50

% Block Depth

Late

2.5

0.6 0.5 0.4 0.3

2.0 Recurrence

0.6

0.015

Llama 3.2 1B

0.7 Residual Entropy

0.8

Coda Colsum Concentration

Mixing Score

Sink Rate

Prelude

0.020

1.0

1.5 1.0 0.5 0.0

100

0

50

% Block Depth

100

0

50

% Block Depth

100

Early

% Block Depth

Figure 52. Stages of inference for retrofitted Llama with each of 128 recurrences, visualized over percentage recurrent depth. These quickly reach a fixed point and do not deviate far from their starting stages of inference.

0.4 0.2

0.020

0.015

0.010

0.0 0

20

40

60

80

% Recurrent Depth

100

0

20

40

60

80

% Recurrent Depth

100

0.5 Recurrence

0.6

Colsum Concentration

Mixing Score

Coda Prelude OLMo2

0.8 Sink Rate

Late

0.6

1.0

0.4 0.3 0.2 0.1 0

20

40

60

80

100

Early

% Recurrent Depth

Figure 53. Stages of inference for retrofitted OLMo-2 with each of 128 recurrences, visualized over percentage recurrent depth. These quickly reach a fixed point and do not deviate far from their starting stages of inference.

35

A Mechanistic Analysis of Looped Language Models

E.6. How Architecture Choices Affect the Formation of Stages of Inference Here we present additional results to supplement those in Sec. 5.1. The extended stages of inference metrics for the models visualized in Fig. 10 are presented in Figs. 54 to 56. Equivalent models with added input injection are visualized in Figs. 57 to 59, and equivalent models without sandwich layers in Figs. 60 to 62.

0.75 0.50 0.25

0.05

0.04

0.03

0.00 0.0

0.2

0.4

0.6

0.8

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

Late

0.4

Recurrence

Prelude Coda Feedforward

Mixing Score

Sink Rate

1.00

Colsum Concentration

It appears in these small scale experiments that feedforward stages of inference are most closely replicated without input injection, and using sandwich layers, as in Figs. 54 to 56. The addition of input injection appears to mean that the final recurrence follows the feedforward stages of inference more closely, but to the detriment of earlier recurrences.

0.3 0.2 0.1

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 54. Stages of inference metrics for a small-scale Looped Transformer of configuration (2, 4 ⊗ 4, 2), compared to a “control” feedforward Transformer with the same training configuration and depth 12.

0.50 0.25

0.05 0.04 0.03

0.00 0.0

0.2

0.4

0.6

0.8

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

0.4 Recurrence

0.75

Mixing Score

Sink Rate

Prelude Coda Feedforward

Colsum Concentration

Late

1.00

0.3 0.2 0.1

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

0.50 0.25 0.00

0.05 0.04 0.03 0.02

0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

Late

0.5 0.4

Recurrence

Sink Rate

0.75

Mixing Score

Prelude Coda Feedforward

1.00

Colsum Concentration

Figure 55. Stages of inference metrics for a small-scale Looped Transformer of configuration (2, 8 ⊗ 4, 2), compared to a “control” feedforward Transformer with the same training configuration and depth 12.

0.3 0.2 0.1 0.0

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 56. Stages of inference metrics for a small-scale Looped Transformer of configuration (2, 12 ⊗ 4, 2), compared to a “control” feedforward Transformer with the same training configuration and depth 12.

36

A Mechanistic Analysis of Looped Language Models

0.50 0.25 0.00

0.05 0.04 0.03 0.02

0.0

0.2

0.4

0.6

0.8

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

Late

0.4 0.3

Recurrence

0.75

Colsum Concentration

0.06

Prelude Coda Feedforward

Mixing Score

Sink Rate

1.00

0.2 0.1

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 57. Stages of inference metrics for a small-scale Looped Transformer of configuration (2, 4 ⊗ 4, 2) 𝐼 , compared to a “control” feedforward Transformer with the same training configuration and depth 12.

0.50 0.25 0.00

0.05 0.04 0.03 0.02

0.0

0.2

0.4

0.6

0.8

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

Late

0.4 0.3

Recurrence

0.75

Colsum Concentration

0.06

Prelude Coda Feedforward

Mixing Score

Sink Rate

1.00

0.2 0.1

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

0.75

0.05

0.50 0.25

0.04 0.03 0.02

0.00 0.0

0.2

0.4

0.6

0.8

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

Late

0.4 0.3

Recurrence

Prelude Coda Feedforward

Mixing Score

Sink Rate

1.00

Colsum Concentration

Figure 58. Stages of inference metrics for a small-scale Looped Transformer of configuration (2, 8 ⊗ 4, 2) 𝐼 , compared to a “control” feedforward Transformer with the same training configuration and depth 12.

0.2 0.1

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 59. Stages of inference metrics for a small-scale Looped Transformer of configuration (2, 12 ⊗ 4, 2) 𝐼 , compared to a “control” feedforward Transformer with the same training configuration and depth 12.

0.50 0.25

0.05

0.04

0.03

0.00 0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

Late

0.4

Recurrence

0.75

Colsum Concentration

Feedforward Mixing Score

Sink Rate

1.00

0.3 0.2 0.1 0.0

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 60. Stages of inference metrics for a small-scale Looped Transformer of configuration (0, 4 ⊗ 4, 0), compared to a “control” feedforward Transformer with the same training configuration and depth 12.

37

A Mechanistic Analysis of Looped Language Models

0.50 0.25

0.05 0.04 0.03

0.00 0.0

0.2

0.4

0.6

0.8

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

Late

0.5 0.4

Recurrence

0.75

Colsum Concentration

Feedforward Mixing Score

Sink Rate

1.00

0.3 0.2 0.1

1.0

0.0

Percentage Depth Through Recurrent Block

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 61. Stages of inference metrics for a small-scale Looped Transformer of configuration (0, 8 ⊗ 4, 0), compared to a “control” feedforward Transformer with the same training configuration and depth 12.

Late

0.50 0.25

0.05 0.04 0.03

0.00 0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

0.0

0.2

0.4

0.6

0.8

1.0

Percentage Depth Through Recurrent Block

0.4 Recurrence

Mixing Score

Sink Rate

0.75

Colsum Concentration

Feedforward

1.00

0.3 0.2 0.1 0.0

0.2

0.4

0.6

0.8

1.0

Early

Percentage Depth Through Recurrent Block

Figure 62. Stages of inference metrics for a small-scale Looped Transformer of configuration (0, 12 ⊗ 4, 0), compared to a “control” feedforward Transformer with the same training configuration and depth 12.

38

A Mechanistic Analysis of Looped Language Models

F. Looped Floorplan To illustrate the “similar attention patterns between recurrences” that we have discussed throughout the paper, in Fig. 63 we visualize all attention patterns for the Retrofitted Llama model on the test prompt. Increasing depth in the model is aligned with increased height up the page; Prelude and coda are outlined in blue and red respectively and separate recurrences are separated by a space.

Figure 63. Entire attention pattern floorplan for the retrofitted Llama model, illustrating the cyclic similarity between recurrences.

39

Record · ID 10332 · SHA-256 d9739a3d53b7e94f
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.